Perspective no_lock Open Access

Estimating the Importance of Change in Patient-Reported Outcome Measures

  • Ron D. Hays
  • Steven P. Reise

Submitted: Apr 27, 2026| Posted: Jun 1, 2026| Published: Jun 1, 2026 | DOI: https://doi.org/10.70542/rcj-japh-art-rlxacc

100

1.4k

 

 

 

Citations

Views

Downloads

Comments

Views

100

1.4k

 

 

 

Citations

Views

Downloads

Comments

Views

search_icon
search_icon Abstract
search_icon Introduction
search_icon Terminology related to important score differences
search_icon Approaches to Estimating Important Change
search_icon Discussion
search_icon Inappropriate Uses of Important Change Estimates
search_icon Future Directions
search_icon References
Abstract
Introduction
Terminology related to important score differences
Approaches to Estimating Important Change
Discussion
Inappropriate Uses of Important Change Estimates
Future Directions
References
Article search_icon Tools search_icon
search_icon
search_icon Abstract
search_icon Introduction
search_icon Terminology related to important score differences
search_icon Approaches to Estimating Important Change
search_icon Discussion
search_icon Inappropriate Uses of Important Change Estimates
search_icon Future Directions
search_icon References
Peer Reviews
Authors
Reader Comments
Supplemental Materials

Abstract

Patient-reported outcome measures (PROMs) are essential for capturing the patient’s perspective on their functioning and well-being in physical, mental, and social health domains. However, statistically significant differences in change can be trivial in magnitude. This commentary reviews the evolving terminology and methodologies for estimating the importance of change in PROMs. We summarize instrument-defined health transitions, bookmarking, and anchor-based estimation methods: mean change, regression, receiver operating characteristic curves, predictive modeling, adjusted predictive modeling, longitudinal confirmatory factor analysis, and longitudinal item response theory. We argue against using distribution-based indices (e.g., SD) as proxies for the importance of change estimates and caution against using group-level thresholds to identify individual change or treatment responders. Regulatory bodies should move toward a unified taxonomy and standardized terminology. In addition, payers must ensure that the thresholds used to indicate important change are robust. Future research needs to prioritize the use of multiple anchors and the integration of statistically significant individual change with patient-perceived change to better categorize treatment outcomes. Researchers should be required to assess person fit and explain how they handled inconsistent data in their analyses.

Keywords: minimally important difference, meaningful score difference; meaningful within-individual change; minimally important change; treatment responder; patient-reported outcomes

Introduction

Patient-reported outcome measures (PROMs), along with observer-reported, clinician-reported, and performance-based measures, are clinical outcome assessments.1  PROMs assess functioning and well-being regarding physical, mental, and social health using survey instruments.2  The number of questions in PROMs varies from a single item (e.g., In general, how would you rate your health?)3 to more than 100; for example, the current Patient-Reported Outcomes System (PROMIS®) physical function item bank contains 165 items and is usually administered using computer-adaptive testing.4  While PROMs are ideally self-administered to capture the patient’s unique perspective, proxy reports (e.g., from spouses or caregivers) have been used for young children and individuals with cognitive impairments.5 Proxy-patient agreement is stronger for observable physical functioning than for internal emotional states such as anxiety.6,7,8

Information about the reliability and validity of PROM measures is essential for interpreting scores at a single point in time and score changes.9  Reliability is the extent to which a PROM yields the same score for a person when no underlying change has occurred. It can be estimated at a single time point (internal consistency) or longitudinally across two or more time points (test-retest).10 Single items with all response options labeled (e.g., poor, fair, good, very good, excellent) are potentially easier to interpret than multi-item scales, but scales are more reliable than single-item scales if the correlations among the items are positive.11

Reliability is a prerequisite for validity, but it does not guarantee it. Validity is the extent to which an instrument measures the intended concept, such as physical functioning, rather than something else. Validity consists of “logical evidence regarding the content or construct being measured, and empirical verification of intended relationships among measures and actions in the world” (p. 1).12 Construct validity is assessed by evaluating whether the direction and magnitude of associations between a PROM score and other variables and is evaluated by examining the extent to which PROM scores at a single time point, or change in scores (responsiveness to change), are associated with a priori hypotheses.13 Validity and interpretation of scores depend on the purpose.12 

Trivial between-group differences in PROM scores can be statistically significant in large-sample studies.  Researchers have spent the past four decades developing methods and identifying differences that are important to patients and other stakeholders.14  The following section reviews the evolution of terminology and approaches for estimating important score differences.

Terminology related to important score differences

There is an alphabet soup of terminology for important score differences (see Table 1). The minimal clinically important difference (MCID) and minimally important difference (MID) refer to the smallest difference in a PROM score that patients perceive as beneficial.14,15 Because the MCID and MID have been used to refer to both between-group and within-group differences in mean scores, minimal important change (MIC) was introduced to represent, putatively, important within-individual change.16 

The United States Food and Drug Administration (FDA) defined a “responder” as a meaningful within-person change specific to the clinical trial population (p. 37)17 and a meaningful score difference as a clinically meaningful within-patient change.18  Meaningful within-person change has been defined as: the amount of change for an individual that indicates a relevant treatment benefit or worsening.19 Authors of a working group of the Drug Information Association described meaningful within-patient change as the “proportion of people who benefit from treatment” (p. 341).19 A recent survey of 248 healthcare professionals, industry, and regulatory stakeholders revealed that 83% felt meaningful change was not well defined and that there was no consensus on the most appropriate methods for estimating it.19  Similarly, it was noted that a meaningful treatment effect “is alarmingly vague” (p. 1885).20

 

Approaches to Estimating Important Change

The patient's acceptable symptomatic state assesses whether the patient is satisfied with the treatment results.  The maximal outcome improvement refers to: (Improvement in score) / (Maximum possible improvement from baseline).  Substantial clinical benefit is indicated by the patient’s perception of substantial improvement.21 The instrument-defined health transitions approach was proposed for preference-based PROMs such as the EQ-5D, SF-6D, and the Health Utilities Index.22 The method assumes that a change to a single level of a single attribute (the smallest health transition) is minimally important. The mean or median of the differences in preference-based scores between health states that differ by only one attribute constitutes the estimate.  Because there is no anchor, the transitions used may be perceived by patients as unimportant. The “idio scale judgment” method asks patients to judge whether hypothetical vignette states are worse, the same, or better than their own state to identify the difference that would make a difference in their daily lives.23  Many methods for estimating important changes (described below) require external information (e.g., retrospective measures of change or clinical parameters) as anchors.  The best anchors are correlated with the change score at 0.371 or higher24 and identify a spectrum of change from substantial improvement to substantial decrements, including those who have changed but not too much—that is, those deemed to have a “minimally important” change. Evaluating the MIC is a special case of responsiveness to change.25  Retrospective change items are often used as anchors.  For example, one might ask patients: In general, how is your physical functioning now compared to before you had COVID-19?  Much better, A little better, About the same, A little worse, Much worse.  

There are at least five methods of estimating important changes that use anchors:

1. Mean change method: Focuses on the average change among those deemed to have “minimal” improvement (e.g., a little better) or decline (e.g., a little worse) on an anchor.

2. Regression methods: These are variants of the mean change method.  In linear regression, change on the PROM is regressed on the anchor (e.g., global rating of change), and the slope (representing the amount of change in the PROM score associated with a unit change on the anchor) is adopted as the estimate.  In mixed-effects regression, the change in the PROM can be modeled across multiple time points, accounting for fixed and random effects.

3. Receiver operating characteristic (ROC) curve method: Identifies the optimal cut point to balance sensitivity and specificity using logistic regression of improved versus not improved for change on the target measure.

4. Predictive Modeling and Adjusted Predictive Modeling: This method finds the change score where the likelihood ratio of change is 1.0 using logistic regression of improved versus not improved for change on the target measure. Adjusted predictive modeling corrects estimates for the proportion of improvement in the sample, the anchor's reliability, and present-state bias (i.e., retrospective ratings can be more strongly influenced by how one feels now than by the difference between one’s current and previous states).

5. Longitudinal confirmatory factor analysis and longitudinal item response theory: Includes an anchor item or items along with prospective PROM measures in factor analysis or item response theory models.

The mean change and regression methods differ from the others in their focus on those who have “minimally” changed (e.g., a little better) relative to the anchor.  The remaining methods dichotomize the anchor (e.g., "a lot better" or "a little better" versus "not better"). Some have suggested that the “traditional focus has been in minimal change…this is now changing…to distinguish between people who had a large improvement/worsening, and a small improvement/worsening” (p. 341).19  Others have argued that “it is best to … leave the ‘M’ to case-by-case, context-based interpretation” (p. 896).26  

When using mean change or regression methods, it is valuable to assess whether there is a monotonic relationship between anchor levels and changes in the PROM score. Ideally, the largest positive change will be seen for those with the most positive category of the anchor (e.g., a lot better), the largest negative change for those with the most negative category of the anchor (e.g., a lot worse), and a change of about zero for those classified as about the same on the anchor. In a study with a treatment and control group, it is informative to compare the change in the target PROM for those deemed to have changed, relative to an anchor, with the change in the control group. If the changes for the two groups are similar, the estimate is potentially problematic unless there is an explanation for why the control group change exceeded the important change threshold. If the important change threshold is 5 and the control group change was 3, the change of 3 does not constitute an important change.  

A simulation found that estimates from the mean change method are more biased than those from other methods.27 Terluin et al.28 concluded that the predictive modeling approach was more precise than the ROC estimation method. Because predictive modeling underestimates the important change threshold when fewer than half of the sample have improved, and overestimates it when more than half have improved, adjusted predictive modeling was recommended.29 But these adjusted predictive modeling estimates were “also biased by the proportion improved” (p. 1274).30 Moreover, authors of a recent simulation study suggested that adjusted predictive modeling is appropriate only when present-state bias is less than 40%.31  An improved adjusted predictive modeling formula was proposed to account for the reliability of transition ratings in addition to the proportion of improved patients.32

The latest estimation approaches are longitudinal confirmatory factor analysis (LCFA) and longitudinal response theory (LIRT). LIRT and LCFA use an anchor item (e.g., retrospective rating of change) as an indicator of latent change in a longitudinal multidimensional IRT and confirmatory factor models, respectively.  The anchor item is allowed to load on both the time 1 and 2 latent variables.30,33 The important change estimate is the threshold (location) estimate for the category of the transition item representing minimally important change (e.g., a little better).

A critical assumption is that the anchor is a valid indicator of latent change in the construct assessed by the PROM.  The LCFA and LIRT methods require that the data fit the statistical model and that the anchor exhibits minimal present-state bias (retrospective ratings of change are more influenced by the present state than by the prior state) or that the bias is estimated.31 The authors of these methods suggest that they are preferable to the other methods reviewed above because they account for measurement errors in both the change score and the anchor item. These new approaches are promising, but much more work is needed to assess their practicality and robustness in real-world applications.

Discussion

Efforts to define and estimate important differences in PROMs have evolved over time.  The current perspective reviews the varying definitions and methods used to assess whether group-level changes in research are large enough to be considered important. Conclusions about the relative performance of different estimation methods in prior work have relied largely on simulations. But the selected simulation variables do not fully capture the reality of real-world data.  Simulations can provide information about the potential effects of the percentage of the sample that improves and the strength of the correlation between the anchor and the change in the target measure on precision, but they cannot estimate what constitutes an important change for patients or other stakeholders. As a result, no single estimation approach can be recommended as superior to the others based solely on simulations.

Interpretation of scores depends on the purpose.12 “Fit for purpose” means that a PROM may be appropriate for one purpose but not another.34  The “purpose” in submissions to the U.S. FDA is labeling and promotional claims about treatment benefits.35  For example, a substantial clinical benefit threshold of change of 22 (1.6 effect size) on the HOOS, JR., and 20 (1.5 effect size) on the KOOS, JR. is being used in the U.S. Centers for Medicare & Medicaid Services mandatory inpatient quality reporting total hip and knee replacement PROMs performance measure. These thresholds were derived from a study where substantial clinical benefit was defined by the difference in mean change for those who reported either “more improvement than I ever dreamed possible” or “great improvement” on a Hospital for Special Surgery satisfaction item: How much did your hip surgery improve the quality of your life?.36

Inappropriate Uses of Important Change Estimates

Investigators have “estimated” the importance of change using distribution-based methods, such as reporting the score associated with 0.5 SD, 0.3 SD, or 0.2 SD. But distribution-based indices are arbitrary constants, not important change estimates.24  When researchers treat these as estimates, it biases important change thresholds towards arbitrary SDs. Instead, actual estimates should be compared with the magnitude of prior estimates to assess their plausibility.  For example, a systematic review found estimates for the EQ-5D-3L, a preference-based measure anchored at 0 (dead) and 1 (perfect health), ranging from 0.003 (minuscule) to 0.952 (extremely large).37 Importance of change estimates should be reported as raw scores and as effect sizes (estimate/SD) to compare with the magnitudes of prior estimates and to interpret them in accordance with existing guidelines.38  If the estimates are implausibly small (e.g., 0.003 for the EQ-5D-3L) or large (e.g., 0.952 for the EQ-5D-3L), they can be excluded.39

Estimates of important changes have sometimes been used to estimate the number of individuals who have achieved an important improvement on PROMs.40,41 But these thresholds are not appropriate for identifying individual change or treatment responders.42,43,44 Responders need to be identified by assessing the significance of individual change.45,46 The minimum detectable change, the smallest amount of change that can be detected beyond measurement error, can be estimated.47,48 In addition, person fit can be evaluated using indices that reflect the consistency between an individual’s responses and the underlying model.49 A person who reports walking 5 miles without difficulty but has some difficulty getting out of bed could be flagged as having a misfit on a physical functioning scale and removed from the analysis.50

Future Directions

Future studies should routinely report both statistically significant individual change and perceived change (e.g., a lot better, a little better, same, a little worse, a lot worse) to enable a four-category classification of individuals:42,43

1. improved significantly and reported at follow-up that they were a little better or a lot better than at baseline.

2. improved significantly but did not report that they were better.

3. did not improve significantly but felt they got better; and

4. neither improved significantly nor reported being better.

Because single items tend to have low reliability,30 multiple anchors should be used in future studies.42

Regulatory bodies should move toward a unified taxonomy and standardized terminology. In addition, payers must ensure that the thresholds used to indicate important change are statistically robust. Future research should prioritize using multiple anchors and integrating statistically significant individual change with patient-perceived change to better categorize treatment outcomes. Researchers should report person-fit indices and explain how they addressed inconsistent data responses in their final analysis.

References

1.

Walton MK, Powers JH 3rd, Hobart J, et al. Clinical outcome assessments: conceptual foundation-report of the ISPOR Clinical Outcomes Assessment - Emerging Good Practices for Outcomes Research Task Force. Value Health. 2015 Sep;18(6):741-752. doi: 10.1016/j.jval.2015.08.006.

2.

Hays RD, Quigley DD. A perspective on the use of patient-reported experience and patient-reported outcome measures in ambulatory healthcare. Expert Rev Pharmacoecon Outcomes Res. 2025;25(4):441-449. doi: 10.1080/14737167.2025.2451749.

3.

Hays RD, Spritzer KL, Thompson WW, Cella D. U.S. general population estimate for "excellent" to "poor" self-rated health item. J Gen Intern Med. 2015 Oct;30(10):1511-1516. doi: 10.1007/s11606-015-3290-x.

4.

Schalet BD, Kallen MA, Perry LM, Garcia SF, Cella D. Putting CATs and item banks to work: how to construct predictive and sensitive PROMIS screeners for use in ambulatory oncology. Qual Life Res. 2025 Oct;34(10):2775-2785. doi: 10.1007/s11136-025-04015-9.

5.

Reimer C, Ali-Thompson S, Althawadi R, O'Brien N, Hickey A, Moran CN. Reliability of proxy reports on patient reported outcomes measures in stroke: an updated systematic review. J Stroke Cerebrovasc Dis. 2024 Jun;33(6):107700. doi: 10.1016/j.jstrokecerebrovasdis.2024.107700.

6.

Hays RD, Vickrey B, Hermann B, Perrine K, Cramer J, Meador K, Spritzer K, Devinsky O. agreement between self reports and proxy reports of quality of life in epilepsy patients. Qual Life Res. 1995;4(2):159-168.

7.

Kroenke K, Stump TE, Monahan PO. Agreement between older adult patient and caregiver proxy symptom reports. J Patient Rep Outcomes. 2022 May 14;6(1):50. doi: 10.1186/s41687-022-00457-8.

8.

Mack JW, McFatrich M, Withycombe JS, Maurer SH, Jacobs SS, Lin L, Lucas NR, Baker JN, Mann CM, Sung L, Tomlinson D, Hinds PS, Reeve BB. Agreement between child self-report and caregiver-proxy report for symptoms and functioning of children undergoing cancer treatment. JAMA Pediatr. 2020 Nov 1;174(11):e202861. doi: 10.1001/jamapediatrics.2020.2861.

9.

Hays RD, Reeve BB. Measurement and modeling of health-related quality of life. In: Quah SR, editor. International encyclopedia of public health. 3rd ed. Academic Press: Elsevier; 2024. p. 352-364.

10.

Hays RD, Edelen MO. Classical test theory. In: Reference module in social sciences. Elsevier; 2025. doi: 10.1016/B978-0-443-26629-4.00122-2.

11.

Johnston E, Reise SP, Spritzer KL, Hays RD. Seeing the light in self-reported glare. Eur J Psychol Assess. 2025;41(3):191-198. doi: 10.1027/1015-5759/a000798.

12.

Shepard LA. Validity for what purpose? Teach Coll Rec. 2013 Sep;115(9):1-12. doi: 10.1177/016146811311500907.

13.

Hays RD, Hadorn D. Responsiveness to change: an aspect of validity, not a separate dimension. Qual Life Res. 1992 Feb;1(1):73-75. doi: 10.1007/BF00435438.

14.

Jaeschke R, Singer J, Guyatt GH. Measurement of health status: ascertaining the minimal clinically important difference. Control Clin Trials. 1989 Dec;10(4):407-415. doi: 10.1016/0197-2456(89)90005-6.

15.

Guyatt GH, Osoba D, Wu AW, Wyrwich KW, Norman GR, Aaronson N, et al. Methods to explain the clinical significance of health status measures. Mayo Clin Proc. 2002 Apr;77(4):371-383. doi: 10.4065/77.4.371.

16.

Terwee CB, Peipert JD, Chapman R, et al. Minimal important change (MIC): a conceptual clarification and systematic review of MIC estimates of PROMIS measures. Qual Life Res. 2021 Oct;30(10):2729-2754. doi: 10.1007/s11136-021-02925-y.

17.

Food and Drug Administration (US). Guidance for industry: patient-reported outcome measures: use in medical product development to support labeling claims. Silver Spring (MD): FDA; 2009 Dec. 39 p.

18.

Food and Drug Administration (US). Patient-focused drug development: incorporating clinical outcome assessments into endpoints for regulatory decision-making [Internet]. Silver Spring (MD): FDA; 2023. Available from: https://www.fda.gov/media/166830/download

19.

Reaney M, Shih V, Wilson A, et al. A consistent lack of consistency: definitions, evidentiary expectations and potential use of meaningful change data in clinical outcome assessments across stakeholders. results from a DIA working group literature review and survey. Ther Innov Regul Sci. 2025 Mar;59(2):337-348. doi: 10.1007/s43441-024-00739-x.

20.

Weinfurt K. Interpreting the meaningfulness of treatment effects estimated in parallel groups designs: comment on Trigg et al. Qual Life Res. 2025 Jul;34(7):1885-1889. doi: 10.1007/s11136-025-03952-9.

21.

Rossi MJ, Brand JC, Lubowitz JH. Minimally clinically important difference (MCID) is a low bar. Arthroscopy. 2023 Feb;39(2):139-141. doi: 10.1016/j.arthro.2022.11.001.

22.

Luo N, Johnson J, Coons SJ. Using instrument-defined health state transitions to estimate minimally important differences for four preference-based health-related quality of life instruments. Med Care. 2010 Apr;48(4):365-371. doi: 10.1097/MLR.0b013e3181c162a2.

23.

Cook, KF, Kallen MA, Coon CD. et al. Idio Scale Judgment: evaluation of a new method for estimating responder thresholds. Qual Life Res 2017 Nov; 26(11): 2961–2971. https://doi.org/10.1007/s11136-017-1625-2

24.

Hays RD, Farivar SS, Liu H. Approaches and recommendations for estimating minimally important differences for health-related quality of life measures. COPD. 2005 Mar;2(1):63-67. doi: 10.1081/copd-200050663.

25.

Revicki D, Hays RD, Cella D, Sloan J. Recommended methods for determining responsiveness and minimally important differences for patient reported outcomes. J Clin Epidemiol. 2008 Feb;61(2):102-109. doi: 10.1016/j.jclinepi.2007.03.012.

26.

Vickers A, Nolla K, Cella D. Drop the "M": minimally important difference and change are not independent properties of an instrument and cannot be determined as a single value using statistical methods. Value Health. 2025 Jun;28(6):894-897. doi: 10.1016/j.jval.2024.09.018.

27.

Griffiths P, Sims J, Williams A, et al. How strong should my anchor be for estimating group and individual level meaningful change? A simulation study assessing anchor correlation strength and the impact of sample size, distribution of change scores and 1methodology on establishing a true meaningful change threshold. Qual2 Life Res. 2023 May;32(5):1255-1264. doi: 10.1007/s11136-022-03286-w.

28.

Terluin B, Eekhout I, Terwee CB, de Vet HC. Minimal important change (MIC) based on a predictive modeling approach was more precise than MIC based on ROC analysis. J Clin Epidemiol. 2015 Dec;68(12):1388-1396. doi: 10.1016/j.jclinepi.2015.03.015.

29.

Terluin B, Eekhout I, Terwee CB. The anchor-based minimal important change, based on receiver operating characteristic analysis or predictive modeling, may need to be adjusted for the proportion of improved patients. J Clin Epidemiol. 2017 Mar;83:90-100. doi: 10.1016/j.jclinepi.2016.12.015.

30.

Bjorner JB, Terluin B, Trigg A, et al. Establishing thresholds for meaningful within-individual change using longitudinal item response theory. Qual Life Res. 2023 May;32(5):1267-1276. doi: 10.1007/s11136-022-03172-5.

31.

Terluin B, Fromy P, Trigg A, Terwee CB, Bjorner JB. Effect of present state bias on minimal important change estimates: a simulation study. Qual Life Res. 2024 Nov;33(11):2963-2973. doi: 10.1007/s11136-024-03763-4.

32.

Terluin B, Eekhout I, Terwee CB. Improved adjusted minimal important change took reliability of transition ratings into account. J Clin Epidemiol. 2022 Aug;148:48-53. doi: 10.1016/j.jclinepi.2022.04.018

33.

Terluin B, Trigg A, Fromy P, Schuller W, Terwee CB, Bjorner JB. Estimating anchor-based minimal important change using longitudinal confirmatory factor analysis. Qual Life Res. 2024 Apr;33(4):963-973. doi: 10.1007/s11136-023-03577-w.

34.

Stewart AL, Hays RD, Ware JE. Methods of validating MOS health measures. In: Stewart AL, Ware JE, editors. Measuring functioning and well-being: the Medical Outcomes Study approach. Durham (NC): Duke University Press; 1992. p. 309-324.

35.

Edwards MC, Slagle A, Rubright JD, et al. Fit for purpose and modern validity theory in clinical outcomes assessment. Qual Life Res. 2018 Jul;27(7):1711-1720. doi: 10.1007/s11136-017-1644-z.

36.

Lyman S, Lee YY, McLawhorn AS, Islam W, MacLean CH. What are the minimal and substantial improvements in the HOOS and KOOS and JR versions after total joint replacement? Clin Orthop Relat Res. 2018 Dec;476(12):2432-2441. doi: 10.1097/CORR.0000000000000456.

37.

Al Sayah F, Jin X, Short H, McClure NS, Ohinmaa A, Johnson JA. A systematic literature review of important and meaningful differences in the EQ-5D index and visual analog scale scores. Value Health. 2025 Mar;28(3):470-476. doi: 10.1016/j.jval.2024.11.006.

38.

Cohen J. Statistical power analysis for the behavioral sciences. 2nd ed. Hillsdale (NJ): Lawrence Erlbaum Associates; 1988. 567 p.

39.

Yost KJ, Sorensen MV, Hahn EA, et al. Using multiple anchor- and distribution-based estimates to evaluate clinically meaningful change on the Functional Assessment of Cancer Therapy-Biologic Response Modifiers (FACT-BRM) instrument. Value Health. 2005 Mar-Apr;8(2):117-127. doi: 10.1111/j.1524-4733.2005.04011.x.

40.

Abu HO, Saczynski JS, Mehawej J, et al. Clinically meaningful change in quality of life and associated factors among older patients with atrial fibrillation. J Am Heart Assoc. 2020 Sep 15;9(18):e016651. doi: 10.1161/JAHA.120.016651.

41.

So R, Terluin B, Takebayashi Y, et al. The minimal important change of the gambling symptom assessment scale among individuals experiencing gambling problems who underwent self-help internet interventions. Int Gambl Stud. 2025;25(2):291-306. doi: 10.1080/14459795.2025.2481839.

42.

Hays RD, Cella D, Peipert JD, et al. Pitfalls of using minimally important or meaningful differences to categorize individual patients. Adv Patient Rep Outcomes. 2025;1(4). doi: 10.1016/j.apro.2025.100302.

43.

Hays RD, Peipert JD. Between-group minimally important change versus individual treatment responders. Qual Life Res. 2021;30(10):2765-2772. doi: 10.1007/s11136-021-02897-z.

44.

Hays RD, Reise SP. Minimally important change estimates should not be used to determine how many individuals have improved over time. Int Gambl Stud. 2026; 1:87-90. doi: 10.1080/14459795.2025.2608603.

45.

McHorney CA, Tarlov AR. Individual-patient monitoring in clinical practice: are available health status surveys adequate? Qual Life Res. 1995 Aug;4(4):293-307. doi: 10.1007/BF01593882.

46.

McLeod LD, Coon CD, Martin S, Fehnel SE, Hays RD. Interpreting patient-reported outcome results: US FDA guidance and emerging methods. Expert Rev Pharmacoecon Outcomes Res. 2011 Apr;11(2):163-169. doi: 10.1586/erp.11.12.

47.

Hays RD, Brodsky M, Johnston MF, et al. Evaluating the statistical significance of health-related quality of life change in individual patients. Eval Health Prof. 2005 Jun;28(2):160-171. doi: 10.1177/0163278705275339.

48.

Hays RD, Spritzer KL, Reise SP. Using item response theory to identify responders to treatment: examples with the Patient-Reported Outcomes Measurement Information System (PROMIS) physical functioning and emotional distress scales. Psychometrika. 2021;86(3):781-792. doi: 10.1007/s11336-021-09774-1.

49.

Reise SP, Flannery WP. Assessing person-fit on measures of typical performance. Appl Meas Educ. 1996;9(1):9-26. doi: 10.1207/s15324818ame0901_3.

50.

Hays RD. Response 1 to Reeve’s chapter: applying item response theory for questionnaire evaluation. In: Madans J, Miller K, Maitland A, Willis G, editors. Question evaluation methods: contributing to the science of data quality. Hoboken (NJ): Wiley & Sons, Inc.; 2011. p. 125-135.