research-edge-series

#008: A Research Survey Is an Instrument, Not a Form

Why your best questions quietly return the wrong answer — and how to design around it

Research Edge Series · By Vinay Thakur


There is a comfortable fiction in applied research: that a survey is just a list of questions, and that a good question is one a reasonable person can read and answer. It feels intuitive. It is also the single most expensive assumption in the field.

A research survey is not a form. A form collects answers. An instrument is built to uncover, test, and decide — and like any instrument, it can be calibrated or silently broken. The breakage rarely shows up in the data. The spreadsheet looks clean. The charts render. The deck gets approved. The error surfaces later, fused into a decision that no longer remembers where it came from.

This piece is about one specific way instruments break — asking respondents to perform a reasoning task they are cognitively incapable of performing, then treating their substitute answer as the one you asked for — and about the fix, which is more interesting than the problem.


1. Watch what surveys actually ask

Here is a question, only lightly disguised from ones genuinely fielded in brand and purpose research:

“Considering this brand’s commitment to sustainability — how much does its environmental purpose increase your preference for the brand?”

It reads fine. It sounds rigorous. It produces a tidy 1–7 distribution, and a brand-strategy decision gets made on the mean.

Now look at what it actually demands. To answer it honestly, a respondent must:

  1. Isolate a single value — sustainability — from everything else they associate with the brand;
  2. Trace its causal contribution to their own preference, holding all else constant; and
  3. Quantify the magnitude of that isolated effect as a number.

No one can do this. Not for your brand, not for any brand, not for any value. People do not have reliable introspective access to the causal structure of their own preferences. This is among the most durable findings in psychology: in their landmark review, Nisbett and Wilson (1977) showed that people confidently report why they preferred or chose something while frequently being demonstrably wrong — the verbal report is constructed after the fact, not read off the actual cognitive process. This finding has been replicated and extended across choice, preference, and judgment tasks over five decades; it is not a contested curiosity but a foundational result.

Asking the question above is asking the respondent to be the analyst. They will decline the appointment — politely, and invisibly.


2. Two failure modes, both invisible

When a question exceeds what the cognitive system can actually do, the response does not arrive as an error. It arrives as a clean number on a clean scale. Two distinct mechanisms produce this, and both are undetectable in the data.

Mechanisms

How broken questions get answered anyway

QUESTION ASKEDExceeds cognitive capacity(no introspective access)SUBSTITUTIONUnconscious swap —answers easierrelated questionSATISFICINGEffort reduction —picks firstacceptable answerCLEAN NUMBER RETURNEDundetectable in the data

Substitution is unconscious — the mind swaps the hard question for an easy one without awareness. Satisficing is effort-driven — respondents pick the first acceptable answer rather than the optimal one. Both produce valid-looking numbers on a clean scale.

The first mechanism is attribute substitution. The standard account in the cognitive-survey literature is that attitude responses are very often constructed on the spot from whatever material is mentally accessible at that moment. Tourangeau, Rips and Rasinski (2000), in the field’s definitive reference text, model survey responding as four stages — comprehension, retrieval, judgment, and response — and show that the judgment stage is the critical vulnerability: when the required integration is genuinely difficult, the mind substitutes a more accessible evaluation without awareness. Schwarz (1999) summarised two decades of evidence bluntly: self-reports are a function of the question, the context, and the momentarily accessible information — not a clean read-out of a pre-existing private fact.

Kahneman (2011) gave this pattern a memorable name: attribute substitution. When the target attribute (what you asked) is hard to assess, the mind swaps in a heuristic attribute (something related but easier) and answers that instead, with no awareness of the trade. The sustainability question’s hard target — the causal contribution of environmental purpose to my preference — is silently replaced by an easy one the mind can answer in milliseconds: how much do I like this brand? You receive a clean number. You label it purpose-driven preference. It was never about purpose at all.

The second mechanism is satisficing. Krosnick (1991) identified a distinct failure mode that operates through a different route: when survey questions are cognitively effortful, respondents sometimes switch from optimising (finding the genuinely best answer) to satisficing (finding the first defensible one). They may select the first scale point that doesn’t feel obviously wrong, agree with whatever direction is implied in the question (acquiescence bias), or endorse the midpoint to signal indifference rather than to communicate a real attitude. Unlike substitution, which is entirely unconscious, satisficing can be partially deliberate — respondents are aware at some level that they are not working hard — but neither mechanism leaves a fingerprint in the data.

The practical upshot is the same in both cases: a number was returned, but the number describes something other than what you asked. The instrument manufactured an answer to a question you never posed.


3. How respondents actually answer: the four-stage model

To know precisely where instruments break, it helps to understand what they are measuring against.

CASM Framework

The four-stage survey response model

1. ComprehensionParse literal meaning; infer intent2. RetrievalAccess relevant memories and beliefs3. JudgmentIntegrate retrieved material← substitution enters here4. ResponseMap judgment onto the scaleTourangeau, Rips & Rasinski (2000)

The judgment stage is the critical vulnerability. When genuine integration is too difficult, the mind substitutes an accessible evaluation and the response process continues as if nothing happened. Nothing in the output distinguishes stage-3 failure from stage-3 success.

The cognitive aspects of survey methodology (CASM) research programme formalised this model across decades of laboratory and field work. Each stage has its own failure modes:

Comprehension is where interpretation variance enters (Tourangeau & Rasinski, 1988). Different respondents often parse the same question differently — the word “consider” in the sustainability example might mean “think about” to one respondent and “given that you accept” to another. Both give you a number; both numbers mean something different.

Retrieval is highly sensitive to what has been recently activated in memory. A prior question can prime a mental frame that biases what material gets retrieved for every subsequent judgment. This is the mechanism behind question-order effects (Section 4 below).

Judgment is the stage where substitution and satisficing enter. If the required integration — weighing, attributing, tracing causality — exceeds what is cognitively feasible, an alternative, accessible evaluation is substituted. The process continues from this point as if the correct judgment had been made.

Response maps the private judgment onto the visible scale. This stage introduces additional distortions: social desirability (shifting the response toward what seems appropriate to report), scale-end avoidance, and context effects from the physical layout of the scale itself. A five-point scale with no midpoint forces a directional response; a seven-point scale allows a non-committal centre; neither accurately captures uncertainty or ambivalence.

The model matters because it localises the problem. Lexically simple questions can fail at the judgment stage. Technically complete scales can introduce systematic error at the response stage. Improving question wording addresses Stage 1; it does nothing about Stages 3 or 4.


4. This is not an amateur problem — and the evidence has teeth

The temptation is to read this as a story about bad researchers. It is not. The most striking demonstrations come from carefully controlled experiments on ordinary respondents, and the effects are large enough to reverse the sign of a correlation.

Order effects

Same questions, different order — completely different result

ORDER ALife sat. asked firstr = −0.12no relationshipORDER BDating asked firstr = +0.66strong relationshipReplication — marital satisfaction:r = .32 → r = .67Schwarz, Strack & Mai (1991)Strack, Martin & Schwarz (1988)

A single ordering decision moved the correlation between the same two variables from near-zero to strongly positive. If questionnaire architecture can generate a correlation, it can also suppress one. Any "what drives what" conclusion is potentially an artefact of design, not a fact about the market.

Strack, Martin and Schwarz (1988) asked students two questions: general life satisfaction, and dating frequency. When life satisfaction came first, the two were essentially unrelated — a correlation of approximately r = −.12. When the dating question came first, priming that domain into accessibility, the correlation rose to approximately r = .66. Same people, same questions, same scale. Only the order changed. The relationship between the two variables shifted from “no link” to “strong positive.” Schwarz, Strack and Mai (1991) replicated the same pattern for marital satisfaction and general life satisfaction — roughly r = .32 in one order, rising to r = .67 when reversed.

Read those numbers as a practitioner. If a single ordering choice can move a correlation from −.12 to .66, then any conclusion you draw about “what drives what” is potentially an artefact of questionnaire architecture rather than a fact about your market. The respondents were not careless. The instrument manufactured the finding.

The order-effect literature extends well beyond these studies. Question-order effects have been documented for attitudes toward abortion (McFarland, 1981), presidential performance evaluations (Halperin, Schwartz & Trevino, 1996), willingness to accept policy tradeoffs (Zaller & Feldman, 1992), and reported behavioural intentions across multiple product categories. The common mechanism is accessibility: a prior question activates a mental frame, and the activated frame biases how subsequent questions are processed. This is not a design flaw in any specific instrument — it is a structural property of how judgment works under time pressure.

A note on replication. Some specific effect magnitudes from the 1980s and 1990s have been revised in subsequent replication attempts, consistent with broader methodological developments in social psychology. The exact figures of −.12 and .66 should be understood as from the original experimental conditions; they may not reproduce identically across all contexts, populations, and question formulations. What has held up robustly across the literature is the existence and direction of the phenomenon: question order affects attitude reports, and the effect can be large. The conservative practitioner conclusion — treat question order as a design variable with empirically testable effects — is supported even under the most sceptical reading of the evidence.


5. You wanted System 2. You got System 1.

It is worth being precise about the gap. You designed the instrument to capture a considered judgment — a deliberate, integrated evaluation. The respondent supplied a fast, intuitive one, automatically, because that is what the cognitive system does under the time and effort constraints of a standard survey. Both feel like answers. Only one matches the construct you set out to measure, and nothing in the data tells you which one you received.

The Mismatch

What the instrument elicits vs. what the decision needs

WHAT YOU DESIGNED FORConsidered, deliberate evaluationintegrated · effortful · System 2WHAT YOU ACTUALLY RECEIVEDFast, intuitive snap judgmentautomatic · low-effort · System 1both look identical in the spreadsheet

A fast intuitive judgment and a considered deliberate evaluation produce identical-looking numbers on a 1–7 scale. The only way to know which you have is to design for it — the data will never tell you.

This matters because the two types of response predict different things. Richetin, Perugini, Adjali and Hurling (2007) showed that implicit (fast, automatic) measures predict spontaneous behaviour, while explicit (deliberate) measures predict deliberative behaviour — the two are differently valid, each for a different kind of downstream decision. An intuitive brand attitude might accurately predict whether someone picks up a product in a supermarket on impulse. It may not accurately predict whether they seek out the brand’s sustainability report or proactively recommend it.

The problem is not that respondents think fast. Fast thinking is accurate and appropriate for many judgments. The problem is a mismatch between the mode the instrument elicits and the construct the decision requires.


6. “Just ask them to think harder” does not work

The obvious fix — instruct respondents to reflect carefully before answering — is weaker than it looks, for two distinct reasons.

First, the empirical evidence for instruction-induced deliberation is mixed at best. Strack and Hannover (1996) found that “think carefully” instructions can sometimes increase consistency effects rather than reduce them, by prompting respondents to construct a more internally coherent narrative — not a more accurate one. The instruction changes the story people tell about their judgment; it does not change the underlying judgment process.

Second, the clean two-systems dichotomy is itself contested in current cognitive science. Evans and Stanovich (2013) defend a broadly valid distinction, but more recent computational and neuroscientific models characterise cognitive operations on a continuum of effort, speed, and automaticity rather than in two discrete types. “System 1” and “System 2” are productive shorthand, not literal descriptions of separate mental modules. This matters because it means there is no instruction that reliably activates a different module — there are only instrument designs that make different cognitive demands.

The practical conclusion is conservative and robust: you cannot instruct System 2 into existence inside a survey. If the considered judgment matters to your decision, it has to be engineered into the instrument’s architecture, not requested from the respondent. Deliberation needs design, not instruction.


7. The formal name for this is validity

Before discussing the fix, it is worth naming what is at stake in the language of measurement theory.

What the sections above describe is a construct validity failure: the instrument is not measuring the construct it claims to measure. Cronbach and Meehl (1955), in the paper that established construct validity as a cornerstone of psychometrics, argued that a measure’s validity cannot be established by face plausibility alone — it must be empirically demonstrated through the pattern of correlations the measure produces with other variables (convergent validity) and through what it fails to correlate with (discriminant validity). A question that asks respondents to report something they have no introspective access to fails construct validity at the source. You can calculate a Cronbach’s alpha on internally consistent nonsense.

Campbell and Fiske’s (1959) multitrait-multimethod matrix formalised the standard. A valid measure of construct X should:

The sustainability question fails all three tests simultaneously. Because it is anchored to overall brand liking rather than the specific attribute, it will converge artificially with overall preference and fail to discriminate between brands with and without genuine environmental associations. It is not a slightly imprecise measure of purpose-driven preference; it is a clean measure of something else that happens to be nearby.

The good news: the validity framework gives precise language for diagnosing and repairing the problem. The question is no longer “is this a good question?” but “does this question measure the intended construct, reliably and separately from other constructs?” The answer determines the fix.


8. The principle: measure what’s answerable, infer the rest

Here is the constructive core.

Stop demanding effort the respondent cannot supply. Instead, decompose the hard construct into components that a fast, intuitive mind can answer accurately, and reassemble the inference yourself. The integration — the genuinely analytical work — is the researcher’s job, not the respondent’s. The instrument’s job is to collect clean, accessible signals; the analyst’s job is to model the relationship between them.

This is not a workaround. It is what measurement is in every mature empirical field. You do not ask a patient to report their cardiac output; you measure observable variables and compute the quantity you actually want. You do not ask a physicist to report a particle’s momentum; you design a detector that captures the signal the particle actually emits and calculate from there. In each case, the hard inference lives in the analysis, not in the subject’s self-report.

The pattern is four steps:

  1. Identify the target construct — what the decision actually needs.
  2. Identify what a respondent can accurately report about that construct (accessible signals — recognition, attribute fit, direct comparison, observed choice).
  3. Design items that collect those signals cleanly, in forms the respondent can answer in seconds.
  4. Use modelling and analysis to infer the target construct from the assembled signals.

The System 2 work happens in step 4. It belongs there.


9. The same question, rebuilt

Take the broken question from Section 1 and re-engineer it.

Decomposition

One unanswerable question → three answerable signals

BROKEN QUESTIONasks respondent to trace causalitySIGNAL 1Associationrecognitiondoes brand = env?SIGNAL 2Revealedchoice taskconjoint / MaxDiffSIGNAL 3Attributefit ratingsingle attributeANALYST INFERS RELATIONSHIPmodelling · regression · causal inference

Each signal is something a respondent can answer in seconds. The causal inference — does environmental association actually move choice? — is produced by the analyst's model, where it belongs, not by the respondent's introspection.

Do not ask: “Does this brand’s environmental purpose increase your preference?”

Instead, collect three signals the respondent can actually provide:

None of these asks the respondent to trace causality. Each is answerable in seconds. You then do the analytical work: model, across the full sample, whether environmental association actually predicts choice controlling for other drivers. The causal claim is now produced by the design and the analysis — where it belongs — rather than extracted from a respondent who never had access to it.

Huber, Wittink, Fiedler, and Miller (1993) showed that choice-based conjoint outperforms direct importance ratings for predictive validity in multiple product categories. This is not a sampling coincidence. Revealed preference tasks outperform stated preference tasks precisely because they bypass the stage at which substitution and satisficing enter — respondents are not reporting internal weights, they are making choices that reveal them behaviourally.


10. The same discipline, everywhere

The sponsorship and purpose case is one instance of a general rule. The same decomposition logic applies across instrument types.

Force trade-offs instead of asking for weights. People cannot accurately report how much they weight an attribute, but they reveal it cleanly in conjoint or MaxDiff choice tasks. Stated importance questions (“How important is price?”) reliably overstate socially approved criteria and understate price sensitivity. The respondent is not lying; they are performing the attribution task and failing at it, exactly as Nisbett and Wilson (1977) predict. Revealed choice cuts through this because no introspection is required — the weight is inferred from the pattern of choices, not reported.

Use comparison and anchoring, not free-floating abstract scales. A judgment relative to a concrete referent is one the cognitive system can make reliably; an absolute free-floating scale forces construction from scratch. “How does this compare to what you normally use?” is more tractable than “How good is this product?”. The former gives the mind a retrieval target. The absolute version asks for an evaluation it constructs fresh each time, making it highly sensitive to context and order effects.

Rotate question order and pretest every item. Given the −.12 to .66 result, treat question order as a design variable with measurable effects on your findings — not a formatting afterthought. Split-sample order rotation is standard practice for longer instruments. Where rotation is not feasible, the general-to-specific principle (broad attitude questions before specific sub-component questions) is supported by the conversational logic analysis in Schwarz, Strack, and Mai (1991).

Measure close to the moment. Retrospective surveys compress weeks or months of experience into a single judgment formed in minutes. That judgment is a reconstruction, not a retrieval — and reconstructions are systematically biased toward the most recent and most intense experiences (the peak-end rule; Kahneman, Fredrickson, Schreiber, & Redelmeier, 1993). Where timing varies, use event-cued recall prompts, time-reference anchors, and experience sampling where measurement quality is paramount.

Use validated multi-item scales for latent constructs. A latent construct — brand equity, customer satisfaction, trust, perceived quality — does not exist in a single observable item. Validated scales have known psychometric properties, including reliability coefficients, convergent and discriminant validity evidence, and cross-sample stability, that a bespoke single question cannot demonstrate. For decision-stakes research, the investment in validated measurement is recovered in interpretability, comparability over time, and defensibility. Yoo and Donthu (2001) for brand equity, and the American Customer Satisfaction Index methodology (Fornell et al., 1996), are established anchors.


11. Catching the break before you field: cognitive interviewing

All of the above can still fail untested. Pre-field detection exists, and it is more accessible than most teams assume.

The method is cognitive interviewing (sometimes called verbal protocol analysis), developed within the CASM research programme and documented in detail by Willis (2005). The technique asks a small number of respondents — typically 5 to 15, which is sufficient to identify most systematic problems (Guest, Bunce, & Johnson, 2006) — to think aloud while completing the survey. A trained interviewer probes their interpretation and reasoning at each item:

What cognitive interviewing reliably surfaces is substitution in action. Respondents will say things like: “I just answered how much I like the brand overall” when probed on a specific attribute question. They will express confusion at causal framing: “I don’t know how to separate that out.” They will describe anchoring on a previous question. None of this appears in the collected data. It only surfaces when you ask — which is why you ask before fielding.

The practical bar is lower than most teams assume. Five in-depth think-aloud sessions routinely identify three to five items producing systematic substitution or satisficing. The fix cost at that stage — rewording, decomposing, removing an item — is negligible. The fix cost after a 5,000-sample national study is a new study.

If one pre-field investment improves instrument quality more than any other, cognitive interviewing is it. It is the quality control step that catches the break before it becomes a finding.


12. The honest counter-view

Intellectual honesty requires stating the case against.

The strong “self-reports are arbitrary” reading has been pushed back on substantively. Critics — including detailed re-analyses in the replicability literature, and longitudinal stability work by Schimmack and Oishi (2005) — argue that the dramatic order effects from the Strack et al. and Schwarz et al. studies are context-specific rather than representative of survey performance generally. Their core claim: chronically accessible constructs (those the respondent habitually and frequently thinks about) produce genuinely stable reports that are minimally affected by incidental priming. Most of the variance in well-being reports, for example, is stable over time rather than dependent on the most recent question.

This is a genuine limit on the strong version of the argument, and it deserves full acknowledgment. It means that well-designed surveys of frequently considered topics — satisfaction with a regularly used product, attitudes toward salient political issues — may be less susceptible to order and accessibility effects than the original experiments suggest. The strong claim that all survey data is contextually arbitrary is not supported by the full weight of evidence.

But this limit does not rescue the broken question in Section 1. That question fails for the introspection reason — Nisbett and Wilson’s (1977) finding — which is orthogonal to whether the attitude is chronically accessible. You can have highly stable, frequently considered brand preferences and still have no introspective access to their causal structure. Stability and accuracy are separate properties of a measurement. A stably wrong number is still wrong.

The right posture, therefore, is calibration rather than paranoia. Not all surveys are broken. Not all questions produce substitution. Context-insensitive stable constructs, measured with validated multi-item scales, close to the moment of experience, can produce genuinely meaningful and defensible data. Knowing precisely where instruments bend is what lets you design ones that do not.


13. Build the instrument. Then trust the decision.

A good research survey does not ask people to do the researcher’s job. It asks what they can actually answer, and infers the rest with rigour. That is the entire difference between a form and an instrument — and it is the difference between a decision you can defend and a number that merely looked clean.

The checklist, reduced to essentials:

  1. Does each key question ask for something the respondent has introspective access to? If not, decompose it into signals they do.
  2. Is question order a design variable in your instrument? If not, make it one — test it, rotate it, or apply the general-to-specific sequencing principle.
  3. Are you asking for stated attribute importance, or can you elicit revealed preference through a trade-off task? If the former, consider the latter; the predictive validity evidence consistently favours it.
  4. Do your latent constructs have validated scales? If you are using bespoke single items, know precisely what validity evidence you are forfeiting and why the tradeoff is acceptable.
  5. Have you run cognitive interviews? If not, you are discovering problems in the data instead of in the pilot — at a cost that compounds from every decision built on the resulting study.

The final question worth taking back to every live instrument you work on:

Which of your current questions is secretly asking the respondent to be the analyst?


References

Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105.

Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302.

Evans, J. St. B. T., & Stanovich, K. E. (2013). Dual-process theories of higher cognition: Advancing the debate. Perspectives on Psychological Science, 8(3), 223–241.

Fornell, C., Johnson, M. D., Anderson, E. W., Cha, J., & Bryant, B. E. (1996). The American customer satisfaction index: Nature, purpose, and findings. Journal of Marketing, 60(4), 7–18.

Guest, G., Bunce, A., & Johnson, L. (2006). How many interviews are enough? An experiment with data saturation and variability. Field Methods, 18(1), 59–82.

Huber, J., Wittink, D. R., Fiedler, J. A., & Miller, R. (1993). The effectiveness of alternative preference elicitation procedures in predicting choice. Journal of Marketing Research, 30(1), 105–114.

Kahneman, D. (2011). Thinking, Fast and Slow. New York: Farrar, Straus and Giroux.

Kahneman, D., Fredrickson, B. L., Schreiber, C. A., & Redelmeier, D. A. (1993). When more pain is preferred to less: Adding a better end. Psychological Science, 4(6), 401–405.

Krosnick, J. A. (1991). Response strategies for coping with the cognitive demands of attitude measures in surveys. Applied Cognitive Psychology, 5(3), 213–236.

McFarland, S. G. (1981). Effects of question order on survey responses. Public Opinion Quarterly, 45(2), 208–215.

Nisbett, R. E., & Wilson, T. D. (1977). Telling more than we can know: Verbal reports on mental processes. Psychological Review, 84(3), 231–259.

Richetin, J., Perugini, M., Adjali, I., & Hurling, R. (2007). The moderator role of intuitive versus deliberative decision making for the predictive validity of implicit and explicit measures. European Journal of Personality, 21(4), 529–546.

Schimmack, U., & Oishi, S. (2005). The influence of chronically and temporarily accessible information on life satisfaction judgments. Journal of Personality and Social Psychology, 89(3), 395–406.

Schwarz, N. (1999). Self-reports: How the questions shape the answers. American Psychologist, 54(2), 93–105.

Schwarz, N., Strack, F., & Mai, H. P. (1991). Assimilation and contrast effects in part-whole question sequences: A conversational logic analysis. Public Opinion Quarterly, 55(1), 3–23.

Strack, F., & Hannover, B. (1996). Awareness of influence as a precondition for implementing correctional goals. In P. M. Gollwitzer & J. A. Bargh (Eds.), The Psychology of Action (pp. 579–596). New York: Guilford Press.

Strack, F., Martin, L. L., & Schwarz, N. (1988). Priming and communication: Social determinants of information use in judgments of life satisfaction. European Journal of Social Psychology, 18(5), 429–442.

Tourangeau, R., & Rasinski, K. A. (1988). Cognitive processes underlying context effects in attitude measurement. Psychological Bulletin, 103(3), 299–314.

Tourangeau, R., Rips, L. J., & Rasinski, K. (2000). The Psychology of Survey Response. Cambridge: Cambridge University Press.

Willis, G. B. (2005). Cognitive Interviewing: A Tool for Improving Questionnaire Design. Thousand Oaks, CA: Sage.

Yoo, B., & Donthu, N. (2001). Developing and validating a multidimensional consumer-based brand equity scale. Journal of Business Research, 52(1), 1–14.

Zaller, J., & Feldman, S. (1992). A simple theory of the survey response: Answering questions versus revealing preferences. American Journal of Political Science, 36(3), 579–616.


Methodological note: The empirical correlation figures cited in Section 4 (approximately −.12 / .66 and .32 / .67) are as reported in the source literature and secondary reviews. Specific effect magnitudes may vary across replication conditions, respondent populations, and question formulations — consult the primary papers for exact experimental parameters. The CASM four-stage model and the construct validity framework (Cronbach & Meehl, 1955; Campbell & Fiske, 1959) are standard references in the psychometric and survey methodology literature and have been extensively validated over several decades of applied use.