Valid, feasible, relevant: Choosing outcome measures
A measure can be convenient to collect and still answer the wrong question. Before choosing a survey or spreadsheet field, name the change you want to understand. “Improve communication” is a goal. It is not yet a measurable outcome.
SQUIRE asks authors to explain their measures, operational definitions, and approaches to evaluating measurement quality. [1] For planning, organize that work around three questions: Does the measure represent what you mean? Can you collect it consistently? Will its result matter to the project decision?
Use three checks together
| Check | What you need to establish | A warning sign |
|---|---|---|
| Valid | The result supports the interpretation you intend in this setting. | A confidence rating is presented as demonstrated competence. |
| Feasible | You can obtain usable data within available time, access, and staffing. | The data arrive after your project closes. |
| Relevant | The measure connects to the project aim and a decision stakeholders need to make. | The score changes, but no one knows what the change means for practice. |
These are not interchangeable. A well established instrument may be too burdensome for your setting. A convenient electronic field may not represent the outcome you care about. A relevant outcome may take longer to change than your project allows.
Turn a concept into a definition
Hypothetical example. You introduce a clearer appointment instruction sheet. Counting sheets handed out measures delivery. Asking patients to identify the correct location, arrival time, and preparation step measures a specific aspect of understanding.
An operational definition might be: “The percentage of eligible patients who correctly identify all three instruction elements immediately after the explanation.” Define who is eligible, how the questions are asked, what counts as correct, when the check occurs, and who records it. This is a locally developed check, not automatically a validated instrument.
Suppose 18 of 24 assessed patients answer all three correctly. The result is 75% among those assessed. It is not necessarily 75% of every eligible patient, because some eligible patients may not have been assessed. Keep the assessment coverage visible too.
Ask whether the measure has room to move
A measure that is already near its maximum may show little improvement even when something useful changes. This is a ceiling effect. Also consider whether the measure can detect the kind of change your intervention is expected to produce, rather than any change at all.
Try the proposed collection process on a few fictional records or approved pilot cases. Can two collectors apply the definition similarly? Is the burden reasonable? Can you explain what a higher or lower value means? These questions are more useful than selecting an instrument only because another student used it.
Your next decision
Write a short definition for your primary outcome before creating the survey. Include the unit, scoring rule, timing, data source, and meaning of improvement. Review proxies in The proxy trap: When easy data measures the wrong thing, timing in Can your outcome change in time?, and instrument selection in Finding a validated instrument and getting permission.