Analyzing Likert data: Items vs scale scores

A five-option survey response and a score formed from several items are not the same kind of measurement decision. Before asking which test to use, ask what the number represents.

An individual Likert-type item has ordered categories, such as strongly disagree through strongly agree. The categories have an order, but equal numerical spacing is not guaranteed. A multi-item scale combines related items according to a justified scoring structure. The literature distinguishes these situations and includes debate about when numerical summaries and parametric analyses are reasonable. [1]

Start with the item or the score

Hypothetical example. A single question asks, “I feel prepared to explain this workflow.” Responses range from 1 to 5. A separate instrument contains six items intended to measure the same preparedness domain and provides a supported rule for calculating a mean score.

The first question gives one ordered response. The second gives a composite only if its items and scoring actually support that interpretation. Six unrelated questions do not become a meaningful scale simply because you average them.

What you have A useful descriptive starting point What to avoid
One ordered item Counts and percentages for each response category. Treating a small mean difference as self-explanatory.
Supported multi-item score Score summaries that reflect its distribution and manual. Inventing a total that the instrument does not support.
Several distinct domains A separate score for each supported domain. Assuming one overall score is always appropriate.

Follow the scoring instructions

Check which items belong together, which are reverse scored, how missing items are handled, and what a higher score means. Use the same rules before and after the intervention. A coding error in one reversed item can make the resulting change misleading.

A reliability coefficient is useful evidence, but it does not by itself prove that all items belong to one construct. Similarly, a five-category response format does not determine the appropriate analysis independently of the design and intended interpretation. [1]

Keep the distribution visible

Suppose nearly everyone selects the top response before training. The item has little room to record an increase. A follow-up mean may barely change even when participants report that the training was useful. That does not establish improvement; it raises a question about whether the selected measure could detect it.

When you show item distributions, readers can see disagreement, clustering, and ceiling effects that an average may hide. When you analyze a supported scale score, explain its range and direction so that a numerical difference has context.

Your next decision

Identify your primary score before collection and follow its scoring guidance. Do not run separate tests on every item simply to search for a significant result. Discuss item-level, ordinal, or score-based approaches according to the question and design. See Finding a validated instrument and getting permission for instruments and Paired t-test or Wilcoxon? Analyzing pre-post data for linked numerical scores.

Sources

[1] Sullivan, G. M., & Artino, A. R., Jr. (2013). Analyzing and interpreting data from Likert-type scales. Journal of Graduate Medical Education, 5(4), 541 to 542.