Decision map
How many points a scale should have
Preston and Colman tested scales from 2 to 11 points in a study published in Acta Psychologica in 2000:- Two-, three-, and four-point scales performed worse on reliability, validity, and discriminating power.
- The indices rose up to around seven points.
- Test-retest reliability dropped on scales with more than ten categories.
- Respondent preference was highest on the ten-point scale, closely followed by the seven- and nine-point ones.
- 0–10 on one or two global anchors, such as the week’s rating and confidence for the week ahead. It’s the scale the patient already knows and the one that produces the most expressive chart.
- 1–5, linear or emoji, for the rest of the core, where week-to-week stability matters more than preference.
- Single choice whenever the answer has real behavioral categories, because a well-written option carries more clinical information than a rating.
Three scale rules that aren’t worth breaking
1
Label the endpoints
“0 = worst possible sleep, 10 = best possible sleep”. Without a verbal anchor, the patient recalibrates the ruler every week and the series turns into noise.
2
Keep the same direction across every question
Higher is always better. If a question inverts that, such as pain level, rewrite it: instead of “how much pain did you feel”, ask “how was your physical comfort”.
3
Never change the scale mid-follow-up
A question that was 0–10 and became 1–5 destroys the comparability of the history. That’s why LiveClin locks question editing once the questionnaire is saved: the platform is protecting your chart.
Next steps
Write questions
The twelve writing rules that protect data quality.
Ready questions
Tested questions with the mode and scale already defined.
References for this page
References for this page
- Preston, C. C., & Colman, A. M. (2000). Optimal number of response categories in rating scales: reliability, validity, discriminating power, and respondent preferences. Acta Psychologica.
- Lozano, L. M., García-Cueto, E., & Muñiz, J. (2008). Effect of the number of response categories on the reliability and validity of rating scales. Methodology.