01 · The Question
Is It Still the Same Instrument After Researchers Change It?
A paper says the researchers used a validated instrument. Then, several sentences later, you discover that they removed three items, rewrote others, changed a five-point response scale to seven points, shortened the recall period, selected only one subscale, or created a new total score.
At what point does an established instrument become a different measure?
There is no universal number of allowable changes. Some adaptations may have little practical effect, while others can alter the content, structure, reliability, score meaning, or comparability of the measure. The important question is therefore not simply whether the instrument was modified, but whether the evidence supporting the original version still supports the version actually used.
03 · What You Need to Know
Evaluate the Version Participants Actually Encountered
The name of an established scale can create a reassuring sense of continuity. Yet measurement depends on more than the name printed in the Methods section. The items, instructions, response process, scoring rules, administration conditions, and intended interpretation all contribute to what is actually being measured.
FDA guidance on clinical outcome assessments explicitly addresses existing, modified, and newly developed measures and emphasizes whether an assessment is fit for its particular context of use. Similarly, COSMIN recommends selecting instruments on the basis of evidence concerning their relevant measurement properties and intended construct rather than relying on reputation alone.
First, Establish What Was Changed
Do not stop at the phrase “adapted from.” Find out what adaptation actually means.
Researchers may remove items, add new ones, change terminology, alter instructions, shorten or lengthen the recall period, change response categories, use only selected subscales, alter scoring, convert a self-administered questionnaire into an interview, or change the setting in which responses are collected.
These changes are not methodologically equivalent. You need enough information to reconstruct the version participants actually completed.
Changing Items Can Change Construct Coverage
Suppose a 20-item instrument measures four dimensions of a construct. Researchers remove six questions because they seem repetitive or inconvenient. Even if the remaining questions appear sensible, the shortened version may no longer represent all four dimensions equally well.
Content validity is particularly relevant here. COSMIN emphasizes whether instrument content is relevant, comprehensive, and comprehensible for the construct, population, and context. Removing or rewriting items can change any of those features.
Ask what information disappeared when the items were removed. The important issue is not merely how many questions remain but what portions of the construct those questions represent.
Small Wording Changes Are Not Always Small Measurement Changes
Researchers sometimes alter wording to make an instrument easier to understand or more appropriate for their context. This may be entirely reasonable. Still, wording affects how respondents interpret questions.
Changing “During the past four weeks, how often...” to “Generally, how often...” changes the reference period. Replacing “work” with “study” changes the behavioral context. Replacing a technical term with a simpler one may improve comprehensibility but could also subtly change meaning.
The methodological significance depends on whether the revised item continues to represent the intended construct in the intended way.
Changing Response Options Can Change the Measurement
Response categories are part of an instrument, not decorative formatting. Moving from four response categories to five, changing “never” through “always” into agreement categories, removing a neutral response, or altering numeric anchors can change how participants respond and how scores are distributed.
The same applies to visual analogue scales, frequency categories, severity labels, and other response formats.
Consequently, evidence based on the original response structure should not automatically be attributed to a substantially different one.
Changing the Recall Period Changes the Question
A measure asking about pain “today” is not identical to one asking about pain “during the past month.” A questionnaire about exercise during the previous seven days is not measuring exactly the same reporting task when researchers change the period to six months.
Recall periods can influence what information participants retrieve and how much they must summarize. They may also affect susceptibility to recall-related measurement problems.
If researchers change the recall period, look for a substantive justification rather than assuming that the original evidence applies automatically.
Using Only One Subscale May Be Appropriate
Not every departure from a complete instrument is problematic. A multidimensional instrument may contain established subscales that were explicitly designed to be scored and interpreted separately.
If researchers use one such subscale for its intended construct and appropriate evidence exists for that subscale, there may be little reason for concern.
The situation is different when researchers select a convenient handful of items from several subscales and create a new score. That score may no longer correspond to an established construct or scoring model.
Established subscale
A component designed and supported for separate scoring and interpretation.
Researcher-created subset
A selection of items assembled for the current study whose measurement properties may differ from the original scale or its established subscales.
Changing Scoring Can Change What the Number Means
Even if every original item is retained, researchers can materially change a measure by altering how responses are scored.
They might average scores instead of following the published algorithm, reverse-score different items, combine dimensions into a single total, establish new cutoffs, dichotomize continuous scores, or treat an ordinal score as though it had another interpretation.
Always distinguish evidence supporting the instrument's original scoring system from evidence supporting a newly created scoring approach.
Administration Mode Can Matter Too
A measure originally designed for private self-completion may behave differently when administered face-to-face by an interviewer. A paper questionnaire may be transferred to an electronic interface. Researchers may provide assistance that was not part of the original procedure.
Such changes are not necessarily damaging. Their importance depends on whether they alter the response process or introduce systematic differences relevant to the construct.
For sensitive topics, for example, interviewer administration could affect participants' willingness to disclose information and increase susceptibility to social desirability bias.
Do Not Assume That Recalculating Cronbach's Alpha Solves the Problem
After modifying a scale, researchers sometimes report a new internal consistency coefficient as evidence that the adapted version remains “valid.” This is insufficient.
COSMIN distinguishes internal consistency from content validity, structural validity, reliability, measurement error, criterion validity, construct validity, cross-cultural validity, and responsiveness. Different measurement properties require different evidence.
Watch Out
A modified scale can have highly correlated items and still omit an important part of the construct, change its dimensional structure, or support a different interpretation from the original instrument. A reassuring reliability coefficient does not restore evidence that the modification may have disrupted.
Modification Often Interacts With Population and Language
Researchers frequently modify instruments precisely because they are using them somewhere new. They may simplify wording for younger participants, replace culturally specific examples, translate questions, or remove items they consider irrelevant.
Those decisions can be defensible, but they create overlapping questions. Consider whether the existing evidence applies to the current population and whether the language version itself is adequately supported.
The More the Instrument Changes, the More the Evidence Must Follow the New Version
There is no magical percentage at which an instrument becomes “new.” Removing 10% of items could be trivial in one instrument and devastating in another if those items uniquely represent an essential dimension.
Think conceptually rather than numerically. Ask whether the modification could change what is measured, how participants respond, how scores are calculated, or what those scores mean.
This is another reason a validated instrument does not automatically guarantee valid measurement in a later study.