Statistical vs clinical significance

A statistical result can be clear without being important to practice. A potentially useful change can also be estimated too imprecisely to support a firm conclusion. Statistical significance and clinical or practical importance answer different questions.

A p value describes how unusual the observed test result, or a more extreme one, would be under the specified null model and its assumptions. It does not give the probability that the null hypothesis is true. Nor does a small p value measure the size or value of a change. Interpret statistical evidence alongside design, context, and consequences rather than as a stand-alone verdict. [1]

Read three pieces of information together

Information Question it helps answer
Effect estimate How much difference or change was observed?
Confidence interval How precise is the estimate under the analysis assumptions?
Clinical or practical context Would a change of this size matter enough to justify the effort or burden?

An effect estimate can use the original units, such as minutes, score points, or percentage points. These units are often more interpretable to site partners than a standardized effect size alone.

A hypothetical operational decision

Suppose a project estimates a reduction of four minutes in the time staff spend locating referral information. A hypothetical 95% confidence interval for the reduction extends from a one-minute increase to a nine-minute reduction.

The point estimate is encouraging, but the interval remains compatible with some worsening, little change, and a worthwhile reduction. The data do not establish that there is no benefit merely because zero lies within the interval. They also do not establish a useful benefit merely because the estimated reduction is four minutes.

Now suppose the team had identified a five-minute reduction as its local operational target before collection. That target gives the result context, but it is not automatically a published minimal clinically important difference. A clinical importance threshold needs justification for the measure, population, and decision where it is used.

Avoid two common shortcuts

A very small effect can produce a small p value when the data are sufficiently precise. A large but uncertain estimate can produce a larger p value when information is limited. Neither situation can be judged by the p value alone.

Also keep statistical uncertainty separate from bias. A narrow interval does not correct an inappropriate comparison, selective follow-up, or a measure of the wrong construct. A before and after difference remains vulnerable to other changes occurring during the same period.

Your next decision

Discuss what amount of change would matter before analyzing the data. In your report, state the estimate, uncertainty, practical interpretation, and design limitations together. Replace “the project worked” with a description of what changed, how certain you are, and what the next decision should be. See Sample size and power for small DNP projects for limited information and Sustaining the change after your project ends for continuation decisions.

Sources

[1] Wasserstein, R. L., & Lazar, N. A. (2016). The ASA statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129 to 133.