01 · The Question
What Should You Do When You Do Not Recognize the Measure?
You are reading a paper and encounter an instrument you have never seen before. Perhaps it is a questionnaire with an impressive acronym, a researcher-developed scale, a performance task, an observational coding system, a device-derived measure, or an index assembled from several variables.
Not recognizing it tells you almost nothing about its quality. An unfamiliar instrument may be carefully developed and well supported. A familiar one may be poorly suited to the study in front of you.
The useful question is therefore not, “Have I heard of this measure?” It is: “What evidence would justify interpreting the measurements in the way these researchers do?” Answering that question requires a structured examination of the instrument rather than a reputation check.
03 · What You Need to Know
Work From the Construct Outward
When an instrument is unfamiliar, it is tempting to start with statistics. You find a Cronbach's alpha, see that it is high, and feel somewhat reassured. That is often too early.
Before examining coefficients, determine what the instrument is supposed to measure and what its scores mean. Measurement evidence only makes sense in relation to a specified construct, population, context, and intended interpretation.
First, Identify Exactly What the Measure Is Supposed to Represent
Find the construct definition given by the study authors or the instrument developers. A measure labeled “digital readiness,” for example, might assess technical skills, attitudes toward technology, access to devices, confidence, willingness to adopt technology, or some combination of these.
Then inspect the actual content or measurement procedure whenever possible. What questions are asked? What behaviors are observed? What tasks must participants perform? What does the device record? How are observations converted into scores?
This lets you address the more fundamental question of whether the study appears to measure the construct it claims to measure.
Find the Original Instrument Source
If the paper cites the instrument's development or validation study, follow that citation rather than relying entirely on the current paper's summary. The original source may explain why the instrument was created, how items or tasks were generated, who participated in its development, what scores or subscales it produces, and what evidence was gathered to support their interpretation.
If the current paper gives an instrument name but no citation, insufficient description, or no route to determining what the measure contains, that is itself relevant to your appraisal. Readers need enough information to understand and evaluate how important variables were measured.
Examine Whether the Content Matches the Construct
Content validity concerns whether the content of a measurement instrument adequately represents the construct or outcome being measured. COSMIN's framework emphasizes three related considerations: relevance, comprehensiveness, and comprehensibility.
Relevance concerns whether the content belongs to the construct and is appropriate for the target population and context. Comprehensiveness concerns whether important aspects of the construct are missing. Comprehensibility concerns whether the people involved understand the questions, tasks, instructions, or other content as intended.
These questions are useful far beyond questionnaires. A performance task can omit important capabilities. An observational protocol can focus on behaviors that only partially represent a construct. A device can precisely record a physical signal that is only an imperfect proxy for the phenomenon researchers want to discuss.
Understand How the Score Is Produced
Do not treat the final number as self-explanatory. Determine how responses or observations become the score used in the analysis.
For a questionnaire, this may involve summing or averaging items, reverse-scoring particular responses, creating subscales, applying cutoffs, or transforming raw scores. For an observational measure, researchers may code behaviors according to specified rules. For a performance measure, the score may depend on accuracy, speed, task difficulty, or a scoring rubric.
Pay particular attention to what a higher or lower score means. If the researchers divide participants into categories, investigate how those thresholds were established and whether they are appropriate for the intended use.
Look for Several Kinds of Measurement Evidence
There is no single statistic that establishes the quality of an unfamiliar instrument. Which evidence matters depends on the type of measure and the interpretation being made.
| Question to Ask |
Evidence That May Help |
What It Can Tell You |
| Does the content represent the intended construct? |
Instrument-development work, expert evaluation, target-population input, content-validity studies |
Whether important content is relevant, sufficiently represented, and understood as intended |
| Does the proposed structure make sense? |
Factor analysis or other evidence concerning internal structure |
Whether relationships among items or components are consistent with the proposed scoring structure |
| Are measurements sufficiently consistent? |
Internal consistency, test-retest reliability, inter-rater reliability, or other appropriate reliability analyses |
Whether particular sources of measurement variation are acceptably controlled for the intended use |
| Does the measure relate to other variables as theory predicts? |
Convergent, discriminant, known-groups, criterion-related, or hypothesis-testing evidence |
Whether observed relationships support the proposed interpretation |
| Does measurement behave appropriately across groups? |
Measurement invariance, differential item functioning, cross-cultural validity, or related analyses |
Whether scores may be comparable across relevant populations or groups |
Not every instrument requires every analysis in this table. The relevant evidence depends on what is being measured and how the scores are being used. The goal is not to collect statistical badges. It is to determine whether the available evidence addresses the interpretation that matters in the current study.
Do Not Let Reliability Substitute for Validity
A reliability coefficient can answer an important question, but not the whole question. High internal consistency, for example, does not establish that the items cover the intended construct comprehensively or that the score means what researchers say it means.
COSMIN likewise treats content validity, structural validity, internal consistency, reliability, measurement error, criterion validity, hypothesis testing, and cross-cultural validity as distinct measurement properties or domains rather than interchangeable evidence.
Watch Out
If the only justification for an unfamiliar scale is a high Cronbach's alpha reported in the current sample, you still know relatively little about whether its scores support the intended construct interpretation. Consistent items can consistently measure something narrower or different from what researchers intended.
Check Who the Evidence Came From
Suppose an instrument has substantial supporting evidence, but all of it comes from middle-aged clinical patients in one country. The current study uses the instrument with secondary-school students elsewhere. That does not automatically invalidate the measure, but it creates a transfer question that deserves attention.
Ask whether the evidence is relevant to the participants in the current study. Age, culture, education, clinical characteristics, and other population differences can matter, particularly when they affect how people understand or respond to instrument content.
This is why the question of whether an instrument was evaluated in the same or a sufficiently relevant population deserves separate consideration.
Check the Language and Version
Instrument names can create a false impression of continuity. The scale used in the current study may not actually be identical to the version examined in earlier research.
Determine whether it was administered in the same language for which the supporting evidence was established. If translated, look for information about the adaptation process and evidence supporting use of the translated version.
Likewise, determine whether the researchers shortened the measure, removed questions, rewrote items, changed response options, altered instructions, combined subscales, or changed the recall period. A modified instrument may require additional evaluation rather than inheriting every property attributed to the original version.
Separate Evidence About the Instrument From Evidence About This Use
This distinction is especially important. You may find extensive literature supporting an instrument, but the relevant question is whether that evidence supports the particular interpretation being made here.
A measure can perform adequately for one purpose without being established for another. Screening, diagnosis, group comparison, individual decision-making, longitudinal change, and prediction can impose different measurement demands.
Do not ask simply, “Is this a good instrument?” Ask, “Is there adequate evidence for using these measurements this way, with these participants, for this conclusion?”
07 · A Quick Checklist
A Fast Appraisal of an Unfamiliar Measure
When you encounter an unfamiliar measure, check:
Identify the construct the instrument is intended to measure and how that construct is defined.
Inspect the instrument's items, tasks, observations, indicators, or measurement procedure when available.
Find the original development or validation source rather than relying only on the current paper's summary.
Determine how responses or observations are converted into the score used in the analysis.
Look for evidence relevant to content, structure, reliability, relationships with other variables, or other measurement properties required by the intended interpretation.
Compare the population and context supporting the measure with those in the current study.
Verify the language and version actually administered.
Check whether the researchers changed items, scoring, instructions, response options, or other important features.
Judge whether the evidence supports the specific interpretation and use required by the study.