Setting up your project data
A spreadsheet should preserve what happened, not require the analyst to guess. Before entering data, decide what one row represents and how every variable will be recorded.
“One row per person” works for some projects. It is not a universal rule. A dataset can use one row per person, one row per encounter, or one row per person at each measurement time. What matters is choosing the structure deliberately and keeping it consistent.
Use one value per cell
Hypothetical example. For a simple paired education assessment, a wide layout could look like this:
| participant_id | pre_score | post_score | post_status |
|---|---|---|---|
| P001 | 6 | 9 | Completed |
| P002 | 8 | No response | |
| P003 | 7 | 8 | Completed |
Each row is one participant. The blank post_score for P002 does not mean zero. A separate status explains why the value is unavailable.
A less useful layout would put “pre 6 / post 9” in one cell, combine names and scores, merge cells for groups, or use a cell’s color as the only record of a missing value. Those choices make checking and analysis harder.
For repeated visits, a long layout with participant_id, visit_number, and score may be more appropriate. Multiple rows for a person are legitimate in that structure, but they are not independent people. The file and analysis plan must both recognize the repetition.
Write a short data dictionary
| Variable | Definition and coding |
|---|---|
| participant_id | Project code; no participant name in the analysis file. |
| pre_score | Baseline assessment total; allowed range 0 to 10 in this fictional example. |
| post_score | Follow-up total using the same scoring rule; blank when unavailable. |
| post_status | Completed, declined, no response, or not eligible at follow-up. |
Record units, allowable values, timing, score direction, and derivation rules. For a percentage, retain the numerator and denominator rather than only the calculated percentage. For dates, use one unambiguous format where their collection and retention are approved.
Separate privacy from formatting
A neat file is not necessarily a de-identified file. Removing names may leave dates, uncommon characteristics, or linkage information that permit identification. HHS describes formal de-identification approaches; a student-created ID alone does not establish that those standards are met. [1]
Use the systems, access restrictions, and retention schedule your institution approves. Keep any linkage file separate with restricted access. Do not place identifiable records into personal storage, unapproved email, or an external analysis tool simply because it is convenient.
Your next decision
Create a small fictional dataset and its dictionary before collection. Ask your advisor whether the structure supports the planned analysis. Preserve the original export and document cleaning decisions so that changes can be traced. See Pre-post surveys: Same instrument, linked responses for linkage and Planning for dropout and missing data for missingness.