1 Concept and terminology

Surrogate-induced bias refers to systematic distortions in study conclusions caused by the use of proxy variables—measures that are treated as substitutes for a target construct that cannot be directly observed or measured with sufficient accuracy. Because a surrogate may represent only part of the underlying concept, statistical relationships involving the surrogate can differ from relationships that would have been observed if the target variable were available.

A key feature of this bias is that it does not merely add noise. Depending on how the proxy relates to the target, the resulting error can shift effect sizes, change apparent rankings, alter the direction of associations, or produce misleading confidence in findings.

1.1 Target variable vs. surrogate measure

The target variable is the construct researchers aim to explain, predict, or infer. It may be latent (not directly observable), expensive to measure, or difficult to operationalize. The surrogate measure is an observable indicator used instead. While surrogates can be useful when they track the target reasonably well, they inevitably encode assumptions about how closely their values mirror the target.

In practice, surrogates are often chosen because they are available at scale, are less costly, or can be collected retrospectively. The distinction between “what the study wants” and “what the data contain” is central to understanding surrogate-induced bias.

1.2 Types of bias induced by surrogates

Surrogate-induced bias can take multiple forms:

  • Attenuation bias, where imperfect surrogate-target alignment weakens observed associations.
  • Inflation bias, where the surrogate captures additional variation correlated with outcomes, creating exaggerated effects.
  • Sign reversal, where the surrogate’s relationship with the target differs across conditions or groups such that the apparent direction of association changes.
  • Spurious associations, where the surrogate correlates with the outcome through pathways unrelated to the intended target construct.

These outcomes depend on the surrogate’s measurement properties and on the structure of the data-generating process.

1.3 Common contexts where surrogates appear

Surrogates arise in many empirical settings, including:

  • Clinical and screening contexts, where test results stand in for diagnosis or disease severity.
  • Behavioral research, where logged activity is used as a stand-in for attitudes or psychological traits.
  • Epidemiology and social science, where administrative records serve as proxies for exposure levels or participation.
  • Technology and learning analytics, where system events are used to infer user intent or engagement.

In each setting, the surrogate is treated as informative about a more substantive but less directly measured construct.

2 Mechanisms of surrogate-induced bias

Surrogate-induced bias emerges because the proxy is not identical to the target. The bias depends on how the surrogate relates to both the construct and other variables in the study, including conditioning variables used in analysis.

2.1 Measurement error and construct mismatch

A primary mechanism is construct mismatch, meaning the surrogate only partially reflects the intended target. Even if the surrogate and target are correlated, differences in what each variable captures can distort inference.

2.1.1 Systematic vs. random surrogate error

Surrogate errors may be random—adding variability without systematically shifting estimates—or systematic—introducing directional distortion. Systematic error can occur when the surrogate performs differently across groups, time periods, or contexts.

2.1.1.1 Differential error across groups or conditions

When the relationship between surrogate and target varies by subgroup or experimental condition, observed associations can reflect those differences rather than the target construct itself. For example, if a proxy tends to overstate the target in one group but understate it in another, regression adjustments may amplify or reverse true effects.

Differential error is also a concern in longitudinal designs, where measurement processes may change across waves due to protocol drift or instrument updates.

2.2 Selection and conditioning pathways

Even with unbiased measurement at the individual level, bias can arise through the ways variables are selected, missing, or conditioned upon in analysis.

2.2.1 Collider and conditioning effects

Conditioning on variables that are downstream of both the surrogate and the outcome can open pathways that create associations where none exist. These collider mechanisms are particularly relevant when selection into the dataset depends on both proxy and outcome-related features.

For instance, if inclusion depends on a combination of proxy status and outcome likelihood, then the selected sample may exhibit relationships inconsistent with the full population.

2.2.2 Feedback loops between proxy and outcome

In some systems, the surrogate and outcome influence each other through feedback. For example, if measurement of the proxy depends on processes triggered by the outcome (or vice versa), then the surrogate becomes entangled with the outcome pathway. In such cases, the proxy is no longer a stable indicator of the target; it partly becomes an outcome of the same causal forces.

2.3 Modeling and functional-form assumptions

Statistical models require assumptions about how the proxy relates to the outcome and about the functional forms connecting variables.

2.3.1 Proxy–outcome relationship mis-specification

If researchers assume the surrogate affects the outcome in a particular way, or that a linear association adequately represents the proxy’s role, mis-specification can bias estimates. Proxy-target relationships may be nonlinear, threshold-based, or heterogeneous across covariate strata.

A mis-specified model can also occur when the surrogate is interpreted as representing the target with constant reliability, even though reliability varies across conditions.

2.3.2 Overfitting or underfitting with surrogate inputs

When a surrogate is used in predictive modeling, the model may learn patterns that are specific to how the proxy is measured rather than patterns reflecting the underlying construct. Overfitting can make the relationship appear stronger in-sample but weaker out-of-sample.

Conversely, underfitting can obscure true signal by forcing overly simple structures on the proxy’s relationship to the outcome, producing apparent attenuation and reducing sensitivity to meaningful effects.

3 Diagnosing and quantifying bias

Diagnosing surrogate-induced bias requires both substantive reasoning about the construct and empirical evaluation of proxy performance. Quantifying bias is often challenging because the target may be unavailable, but diagnostics and sensitivity analyses can bound plausible distortions.

3.1 Conceptual checks and construct validity

A foundational step is assessing whether the surrogate operationalization plausibly matches the target construct. Conceptual validity includes:

  • Content coverage: whether the proxy reflects most facets of the target.
  • Directionality: whether higher surrogate values consistently imply higher levels of the target (not merely a correlated but different concept).
  • Boundary conditions: contexts in which the surrogate is expected to fail (e.g., measurement differs by environment, subgroup, or time).

These checks help determine whether bias is likely to be attenuation-like, directional, or possibly even sign-changing.

3.2 Statistical diagnostics for proxy performance

Empirical diagnostics evaluate how well the surrogate performs as an indicator.

3.2.1 Predictive accuracy vs. causal adequacy

High predictive accuracy of the surrogate for outcomes does not guarantee causal adequacy. A surrogate may predict outcomes well because it captures downstream correlates or selection processes, rather than the underlying target. Conversely, a proxy can have modest predictive performance yet still serve as a reasonable indicator under a correctly specified measurement model.

Diagnostics therefore distinguish between prediction (association with outcomes) and measurement (association with the target construct).

3.3 Sensitivity analysis

Sensitivity analysis explores how conclusions change under plausible deviations from the assumptions that justify the surrogate’s use.

3.3.1 Bias amplification scenarios

Bias amplification occurs when small inaccuracies in proxy-target alignment translate into large errors in estimated causal or association parameters. Such amplification can happen when models rely heavily on the surrogate, when the surrogate is used in conditioning sets, or when error correlates with other covariates.

A sensitivity framework can vary reliability, differential error rates, or proxy-target residual correlations to examine how estimates behave under alternative assumptions.

3.4 Benchmarking against direct measurements

When direct measurements of the target are available in a subset, researchers can compare proxy-based inferences with target-based estimates. Benchmarking provides practical evidence about the direction and magnitude of distortion.

Even partial validation can be informative when direct measurements are used to estimate measurement error parameters, which then feed into corrected analyses.

4 Mitigation strategies

Mitigation aims to reduce systematic divergence between surrogate and target, prevent problematic analytic choices, and use models that explicitly represent measurement uncertainty.

4.1 Improving surrogate quality

A straightforward route is to strengthen the proxy’s fidelity to the target.

4.1.1 Calibration and re-scaling approaches

Calibration involves mapping surrogate values onto target-relevant scales, using validation datasets or external standards. Re-scaling can correct for systematic differences in average levels or variance, helping align surrogate distributions with the target construct.

This approach is most effective when the calibration relationship remains stable across time, groups, and settings.

4.1.2 Multi-surrogate or ensemble proxies

Instead of relying on a single proxy, combining multiple surrogates can capture different aspects of the target. Ensemble approaches may reduce sensitivity to any one surrogate’s weaknesses, particularly when errors are partially independent across measures.

Model-based fusion (e.g., weighted combinations) can also help represent uncertainty about the target implied by each proxy.

4.2 Study design improvements

Design choices can reduce reliance on assumptions and improve identifiability.

4.2.1 Collecting validation subsamples

Collecting target measures for a subset enables estimation of surrogate reliability and error structure. Validation subsamples can be planned prospectively to ensure representativeness and adequate coverage of relevant strata.

This also supports benchmarking and informs the choice of correction models.

4.2.2 Pre-specifying proxy usage rules

Pre-specifying when and how proxies are used limits analytic flexibility and clarifies assumptions. Rules may include:

  • thresholds for acceptable proxy performance,
  • criteria for which validation estimates apply,
  • plans for handling missingness in proxy variables.

Pre-specification improves transparency and reduces the risk of post-hoc adjustments tailored to observed results.

4.3 Analytical methods

Analytic techniques can account for measurement uncertainty instead of treating the proxy as truth.

4.3.1 Measurement error models

Measurement error models treat the observed surrogate as a noisy version of the latent target. These models often require additional assumptions or auxiliary data, such as repeated measures, validation samples, or instrument-like sources that inform error variance.

Corrected estimates can be obtained when the error structure is sufficiently identified.

4.3.2 Latent variable and factor approaches

When the target is latent, factor models and latent variable approaches can incorporate multiple indicators (including surrogates) to estimate a shared underlying construct. Such methods can separate common variance attributable to the target from variance specific to each proxy.

The success of these approaches depends on reasonable assumptions about the measurement model and the number and quality of indicators.

4.3.3 Causal inference methods with proxies

For causal goals, proxy handling must align with causal structure. Methods that explicitly model measurement processes, address selection bias, or use proxy-aware identification strategies can help recover interpretable effects.

The key requirement is that the assumptions needed for causal identification are stated clearly and checked as far as possible with available data and validation.

5 Reporting and interpretation

Reporting should make surrogate assumptions auditable and help readers evaluate how much uncertainty is introduced.

5.1 Communicating limitations and uncertainty

Authors should describe surrogate properties, including expected reliability, potential systematic error, and any evidence from validation. Uncertainty intervals should reflect measurement-related uncertainty rather than only sampling variation.

When correction methods are used, reporting should include model components that represent measurement error or proxy-target linkage.

5.2 Interpreting effect estimates with proxies

Effect estimates derived from proxies should be interpreted as statements about the target only insofar as the assumptions justify that link. If assumptions are weak, estimates may reflect a blend of target effects and proxy-specific pathways.

Clear language can prevent over-interpretation, especially when proxies are used in causal or mechanistic claims.

5.3 Documenting surrogate selection rationale

Documentation should include why a surrogate was chosen, what aspect of the target it captures, and what aspects it does not. It should also note alternative measures considered, proxy performance metrics, and any constraints that limited calibration or validation.

This rationale supports reproducibility and helps other researchers judge transportability to new contexts.

6 Applications and examples (illustrative)

Illustrative examples clarify how surrogate-induced bias can arise in everyday research practice, without implying any specific real-world claim.

6.1 Using screening tests as proxies for diagnoses

A screening test can correlate with an underlying condition, but it may not perfectly separate diseased from non-diseased individuals. If sensitivity and specificity differ across demographic strata or environments, estimated differences in diagnosis rates may be distorted. Additionally, if the screening process changes after symptoms appear, the proxy may become partially influenced by the outcome process.

6.2 Behavioral logs as proxies for attitudes

Digital traces such as clicks, watch time, or message length are sometimes used to infer underlying attitudes. Yet these logs may reflect situational factors (recommendation algorithms, interface design, or social context) that are not part of the attitude construct. Model-based interpretations can therefore conflate genuine belief differences with changes in exposure or engagement mechanics.

6.3 Administrative records as proxies for exposures

Administrative data may record formal participation, reported use, or recorded encounters, which may lag behind or miss informal exposure pathways. If missingness is correlated with outcome risk—for instance, if individuals most at risk are more likely to appear in records—then proxy-based exposure estimates can be biased. Even when exposure is correctly recorded, administrative categories may aggregate diverse underlying intensities, leading to construct mismatch.

Surrogate-induced bias is closely related to other measurement and inference issues.

7.1 Measurement error and misclassification

Measurement error covers deviations between measured values and true quantities, including misclassification for categorical variables. Surrogate-induced bias often arises as a specialized form of measurement error, where the measured proxy replaces a latent or hard-to-measure target.

7.2 Proxy variables and substitution bias

Proxy variables are generally substitutes used when direct measures are unavailable. Substitution bias refers to distortion created by replacing the intended variable with a proxy, especially in causal frameworks. Surrogate-induced bias can be understood as a broad category that includes substitution bias effects.

7.3 Construct validity and validity threats

Construct validity evaluates whether an operational measure represents the intended construct. Surrogate-induced bias is a validity threat: if the proxy does not reflect the target construct, conclusions about the construct become unreliable.

Construct validity concerns also guide which mitigation strategies are appropriate, such as calibration, validation sampling, or latent variable modeling.

8 Further reading and methodological resources

Further study typically includes measurement error theory, validation study design, and methods for causal inference under proxy and selection problems. Methodological resources may include textbooks and review articles on measurement error correction, latent variable models, and sensitivity analysis frameworks that support transparent interpretation when true target variables are missing or imperfectly observed.