Paired t-test or Wilcoxon? Analyzing pre-post data
When you measure the same people before and after an intervention, the key comparison is each person’s change. Begin by verifying the links and defining the direction of change. If change equals follow-up minus baseline, a positive value means an increase, which may be good or bad depending on the scale.
The paired t test evaluates the mean of these differences. [1] The Wilcoxon signed rank test uses their signs and ranked magnitudes and relies on a symmetry assumption for the usual location-shift interpretation. [2] They are not interchangeable tests selected solely by the number of participants.
Examine the differences
Hypothetical example. Most of 18 participants improve by one or two points on an assessment, while one participant improves by ten. That unusually large change may influence the mean. Check whether it is a valid observation, a coding mistake, or a scoring error before deciding how to analyze the data.
For a paired t test, the normality concern in a small dataset concerns the distribution of differences, not whether baseline and follow-up scores separately look normal. Independence between participants’ differences also matters. [1]
| Consideration | Paired t test | Wilcoxon signed rank test |
|---|---|---|
| Main focus | Mean within-person difference. | Signed ranks of within-person differences. |
| Distribution issue | Severe nonnormality or influential differences can matter, especially in a small dataset. | The usual location interpretation requires a reasonably symmetric difference distribution. |
| Important practical check | Do numerical differences have a meaningful interpretation? | Can differences be meaningfully ranked, and how are zeros and ties handled? |
The signed rank test should not be described simply as comparing the two observed medians. Its usual interpretation as a change in location also depends on the symmetry of the differences. [2]
Do not use a mechanical cutoff
“Fewer than 30 participants means Wilcoxon” is not a sound decision rule. A small dataset can have well behaved differences, while a larger one can contain serious data problems. A normality test alone is also not the entire decision. Review the shape, unusual observations, score properties, and purpose of the comparison.
Repeated scores and no-change values are common with short assessments. They affect the information available to a rank-based analysis. Report the method used for handling ties and zeros when relevant; do not assume that every software implementation handles them identically. [2]
Report the change, not just the test
Present the number of complete pairs, summaries at each time, an appropriate estimate of change, uncertainty when estimable, and the statistical result. Specify the direction of scoring. Avoid writing that the intervention “caused improvement” merely because a paired test was significant.
Your next decision
Prepare a plot or summary of individual differences and review it with your statistical advisor. If score interpretation, asymmetry, clustering, or severe ceiling effects make both options questionable, discuss an alternative rather than forcing a choice. See Analyzing Likert data: Items vs scale scores and Statistical vs clinical significance.
Sources
[2] Hollander, M., Wolfe, D. A., & Chicken, E. (2014). Nonparametric statistical methods (3rd ed.). Wiley.