1 General concept

1.1 Definition

Validation is the process of determining whether a statement, model, method, or result is sound and appropriate for its intended use. It evaluates whether what is produced—such as a measurement, prediction, or representation—corresponds to the reality it aims to capture or to the purpose it is meant to serve.

1.2 Purpose

The central aim of validation is risk reduction: it seeks to confirm that an approach performs acceptably before it is relied upon for decisions, explanations, or further research. By establishing credibility, validation helps prevent the use of outputs that are misleading, unsupported, or incompatible with the scenario in which they will be applied.

1.3 Validation in the scientific method

In the scientific method, validation supports the movement from plausible claims to substantiated conclusions. Evidence is gathered through observation and experiment, and then evaluated for consistency, explanatory adequacy, and agreement with independent information. Validation also includes checking whether conclusions persist when assumptions are modified or when new data are introduced.

1.4 Distinction from verification

Verification generally asks whether something was implemented correctly relative to a specification or internal rules. Validation asks whether that correctly implemented artifact actually meets the real-world or conceptual need it targets. Put differently, verification emphasizes conformance; validation emphasizes fitness for purpose.

2 Types of validation

2.1 Empirical validation

Empirical validation relies on measurements or observations. A claim or model is tested by comparing its outputs with data gathered under relevant conditions. The strength of empirical validation typically depends on the quality, coverage, and representativeness of the data used.

2.2 Theoretical validation

Theoretical validation draws on logic, mathematics, and domain knowledge. It assesses whether an approach is consistent with established principles, derives correct limiting behaviors, or follows known constraints. Theoretical support often complements empirical testing, especially when experiments are difficult or incomplete.

2.3 Statistical validation

Statistical validation examines how well results hold up under uncertainty. It may include goodness-of-fit measures, calibration checks, hypothesis testing, confidence intervals, and evaluation of predictive distributions. This type focuses on whether observed performance is distinguishable from what could occur by chance or sampling variability.

2.4 Construct validation

Construct validation assesses whether a tool or measure truly reflects the underlying concept it intends to represent. For example, a questionnaire may be tested to confirm that scores behave as theory predicts across related factors and groups. This approach is crucial when the target is abstract, such as attitudes, abilities, or latent traits.

2.5 External validation

External validation evaluates performance using data or settings not used during development. It addresses the question of general applicability: whether a model or method works beyond the environment where it was created, including differences in populations, instruments, workflows, or experimental designs.

3 Validation process

3.1 Formulating criteria

Validation begins by defining what “valid” means for the specific application. Criteria can include accuracy thresholds, acceptable error rates, allowable bias, coverage ranges, assumptions to be satisfied, and operational constraints. Clear criteria help ensure that validation is not merely qualitative judgment.

3.2 Collecting evidence

Evidence is gathered from relevant sources, such as experiments, benchmark datasets, expert knowledge, and independent measurements. Good practice favors diverse evidence that probes the approach across conditions likely to matter in deployment.

3.3 Testing predictions

For predictive models, validation commonly includes testing whether forecasts match observations. The testing may involve held-out datasets, future time periods, or scenario-based simulations. For methods that generate explanations or classifications, performance is assessed against ground truth or trusted reference outcomes.

3.4 Comparing with standards

Comparisons are made against standards that function as reference points. These can be physical standards (for instruments), established datasets (for algorithms), or accepted theoretical results (for frameworks). Deviations from standards indicate limits of validity and can guide refinement.

3.5 Assessing reproducibility

Reproducibility evaluates whether the same approach produces consistent results under similar conditions. It may involve repeating experiments with the same protocol, using the same data processing pipeline, or demonstrating that independent teams obtain comparable findings when applying the method correctly.

4 Validation in scientific research

4.1 Hypothesis validation

Hypothesis validation evaluates whether evidence supports or refutes proposed explanations. Researchers test implications derived from the hypothesis and assess whether observed patterns align with predictions while controlling for alternative explanations. Even when results support a hypothesis, validation typically emphasizes the conditions under which the support holds.

4.2 Model validation

Model validation checks whether a model’s structure and parameters generate outputs that reflect the target phenomenon. This can include verifying assumptions, inspecting residuals or error behavior, confirming that the model reproduces key features of data, and testing generalization to new or more challenging cases.

4.3 Instrument validation

Instrument validation concerns whether an apparatus measures as intended. It often includes verifying calibration, evaluating response linearity, checking drift over time, and confirming that performance remains stable under expected environmental and operational conditions.

4.4 Measurement validation

Measurement validation focuses on whether a measurement procedure yields meaningful values for its intended construct and context.

4.4.1 Accuracy

Accuracy describes how close measurements are to a trusted reference or true value. It is influenced by systematic error, calibration quality, and correct handling of units and models linking the measured signal to the quantity of interest.

4.4.2 Precision

Precision refers to the consistency of repeated measurements under unchanged conditions. High precision can occur even when accuracy is poor, so precision is evaluated separately from accuracy.

4.4.3 Reliability

Reliability reflects whether the measurement yields stable results across repeated applications, different operators, or time. It captures practical consistency in real workflows, often using agreement statistics or repeatability assessments.

4.5 Data validation

Data validation ensures that inputs and derived datasets are coherent and trustworthy. It covers checks such as range validation, missing-data handling, consistency across fields, duplicate detection, unit normalization, and validation of data provenance and transformations.

5 Validation methods

5.1 Controlled experiments

Controlled experiments validate by systematically manipulating factors while holding other variables constant. This design helps isolate causal relationships and tests whether an approach responds to changes in relevant inputs as expected.

5.2 Replication studies

Replication studies repeat research or experiments to evaluate whether results reappear. Replication can be direct, varying little from the original protocol, or conceptual, changing methods while targeting the same underlying claim. Both contribute evidence about stability of findings.

5.3 Cross-validation

Cross-validation is a resampling strategy used in predictive modeling. Data are partitioned into training and evaluation subsets multiple times, providing estimates of performance variability. It helps reduce dependence on a single data split, though its reliability depends on appropriate partitioning and assumptions about data independence.

5.4 Peer review

Peer review evaluates research outputs through expert scrutiny of methods, assumptions, analysis choices, and reporting clarity. While peer review does not guarantee correctness, it serves as an external quality filter by exposing weaknesses, ambiguities, and inconsistencies before broad adoption.

5.5 Sensitivity analysis

Sensitivity analysis studies how outputs change when inputs, parameters, or assumptions are varied within plausible bounds. It helps identify whether conclusions depend strongly on uncertain choices and whether the method’s behavior remains stable across reasonable scenarios.

6 Standards and quality control

6.1 Acceptance criteria

Acceptance criteria define thresholds that determine whether an outcome is adequate. These may include maximum error tolerances, minimum statistical power, required coverage, or performance metrics tied to safety, cost, or usability requirements.

6.2 Calibration

Calibration aligns an instrument or system’s responses with known references. Proper calibration reduces systematic deviation and supports traceability, enabling later measurements to be interpreted meaningfully.

6.3 Quality assurance

Quality assurance encompasses procedures that prevent defects and promote consistent execution. It includes documentation standards, protocol adherence, training, audit trails, and systematic monitoring of processes used to generate results.

6.4 Error detection

Error detection identifies problems arising from data recording, computation, instrument behavior, or procedural deviations. Techniques can include automated validation rules, anomaly detection, control samples, and independent verification steps aimed at catching mistakes early.

7 Limitations and challenges

7.1 False validation

False validation occurs when an approach appears to perform well under limited testing but fails in broader or more realistic settings. Causes include unrepresentative data, overconfident metrics, insufficient coverage of conditions, or reliance on shortcuts that do not reflect the full complexity of the target.

7.2 Bias and confounding

Bias arises when evidence systematically skews results toward particular outcomes or groups. Confounding occurs when an unaccounted factor influences both inputs and outputs, making it unclear whether the observed relationship reflects the intended mechanism rather than an artifact of the data.

7.3 Overfitting

Overfitting is a form of poor generalization where a model captures noise or idiosyncrasies of the training dataset. Validation helps reveal this by evaluating performance on data not used for fitting, but it must be implemented with careful partitioning to avoid leakage and misleading estimates.

7.4 Incomplete evidence

Validation depends on evidence availability. If tests cover only a narrow range of conditions, the confirmed validity may not extend to other contexts. Incomplete evidence is especially problematic when new operating conditions introduce changes in data distributions, constraints, or measurement conditions.

7.5 Context dependence

Many validation claims are conditional. Performance can change across environments, populations, instruments, or workflows due to differences in noise levels, missingness patterns, sampling procedures, or underlying mechanisms. As a result, validation statements typically specify the context in which they are expected to hold.

8.1 Verification

Verification checks whether an implementation matches its specified requirements. It is often performed internally against documentation, code logic, or designed procedures rather than against the external reality the artifact is intended to represent.

8.2 Falsification

Falsification is a principle in which claims are evaluated by attempting to find counterevidence. If evidence contradicts predictions or assumptions, the claim is weakened or rejected, thereby contributing to the validation landscape.

8.3 Reliability

Reliability concerns consistency of results across repeats and conditions. High reliability improves confidence in validation outcomes because it reduces uncertainty attributable to random variation.

8.4 Generalizability

Generalizability measures whether an approach’s performance extends beyond the data or conditions used for evaluation. It is closely related to external validation and is often the determining factor for whether results can be applied broadly.

8.5 Reproducibility

Reproducibility focuses on whether independent attempts can obtain consistent results. It strengthens validation by showing that outcomes are not tightly bound to a specific dataset, operator, or hidden implementation detail.