1 What Expert Elicitation Is
1.1 Core purpose and typical outputs
Expert elicitation is a structured procedure for obtaining quantitative or qualitative judgments from individuals with domain expertise when direct measurement is unavailable, incomplete, or prohibitively expensive. The central goal is to translate those judgments into explicit representations of belief—such as point estimates, probability distributions, predictive scenarios, or uncertainty intervals—that can be used in analysis and decision-making.
Typical outputs include quantiles (e.g., median and credible ranges), parameters of probabilistic models, calibrated forecast distributions, and uncertainty summaries that distinguish different sources of uncertainty.
1.2 When it is used in scientific inquiry
In scientific settings, elicitation is often used when evidence is early-stage, sparse, or not yet observable. Common examples include emerging technologies, newly characterized phenomena, experiments with limited samples, or rare-event settings where data accumulation is slow. Elicitation can also support model initialization in situations where direct parameter estimation is not feasible, and it can clarify assumptions when empirical identifiability is limited.
1.3 Key differences from informal expert opinion
Informal expert opinion tends to rely on unstructured discussions that produce narrative conclusions without standardized uncertainty quantification. Expert elicitation differs by using predefined questions, explicit scales, and methods designed to elicit probabilistic information rather than vague judgments. It also typically includes procedures for bias reduction, coherence checks, scoring or incentive mechanisms, and transparent documentation of how the final distribution or interval was formed.
1.4 Types of elicitation targets (estimation, prediction, uncertainty)
Elicitation targets vary by the object of prediction or estimation:
- Estimation focuses on unknown quantities describing the current or static state of the system (e.g., a parameter value).
- Prediction targets future outcomes at specified horizons (e.g., an event probability within a time window).
- Uncertainty elicitation aims to characterize not only the expected value but also the spread, tails, and confidence structure reflecting epistemic uncertainty (lack of knowledge) and, where relevant, other uncertainty components.
2 Elicitation Design and Preparation
2.1 Defining the question precisely
A well-designed elicitation begins with a question that is precise enough for experts to interpret consistently while remaining decision-relevant.
2.1.1 Translating real-world quantities into decision-relevant variables
Real-world concepts must be converted into measurable or definable variables. This translation includes clarifying what counts as an event, how measurements would be taken if possible, and the exact mapping from real processes to the target random variable used in analysis.
2.1.2 Specifying time horizons, units, and scope
Time horizon, units, and scope determine what is being asked and prevent experts from unintentionally responding to a different scenario. The elicitor specifies observation periods, geographic or system boundaries, relevant conditions, and any stratifications needed for the analysis.
2.2 Selecting and recruiting experts
The usefulness of elicitation depends on selecting experts who can supply informed judgments and interpreting their responses in a way that respects independence and expertise quality.
2.2.1 Criteria for subject-matter fit
Experts are chosen based on demonstrated knowledge, experience with comparable problems, familiarity with relevant data-generation processes, and ability to reason about uncertainty. Qualification criteria are ideally documented to reduce arbitrariness.
2.2.2 Managing expertise diversity and independence
Diversity can improve coverage across modeling perspectives, methods, or subdomains. Independence matters because correlated judgments can artificially narrow uncertainty. Preparation therefore often includes identifying overlapping affiliations, shared information sources, or channels through which experts may influence each other.
2.3 Gathering background information and evidence
Before elicitation, the organizer compiles relevant literature, prior studies, data summaries, model assumptions, and known constraints. This material provides a common factual base while still allowing experts to express judgment about parts that are uncertain. Background information is typically provided in a controlled form to avoid leading experts toward a preferred distribution.
2.4 Ethical and practical considerations
Ethical practice includes informed consent, confidentiality where needed, disclosure of incentives or scoring rules, and careful handling of conflicts of interest. Practically, the design should fit available time and expertise, avoid excessive cognitive load, and ensure that the elicitation tools are comprehensible.
3 Elicitation Methods
3.1 Structured interviews and guided judgment
Structured interviews use a standardized script to guide experts through a sequence of questions. The elicitor clarifies terms, checks understanding, and prompts for probabilistic answers where appropriate. While not as formal as some other methods, guided judgment can still improve consistency and support later coherence checks.
3.2 Probability wheel, percentiles, and distribution elicitation
Probability elicitation methods convert beliefs into probability statements. Common approaches include asking for quantiles, using a “probability wheel” that maps odds to probabilities in a visual framework, or collecting enough percentile information to define a full distribution.
3.2.1 Eliciting quantiles vs. moments
Quantiles often provide more robust information about tails than moments such as mean and variance. Moments can be harder to interpret under strong nonlinearity or skewness, and they may inadequately capture behavior in extreme regions. Quantile-based elicitation supports direct reporting of ranges like central credibility intervals.
3.3 Scoring rules and incentive-compatible approaches
Scoring rules reward experts according to the quality of their probabilistic forecasts. Proper scoring rules are designed so that an expert’s best strategy, in principle, is to report their true beliefs. In practice, scoring may be combined with calibration feedback or peer comparison, depending on the setting and feasibility.
3.4 Calibration and “seed” questions
Calibration techniques use questions with known or later verifiable outcomes to assess whether experts systematically over- or under-estimate probabilities. “Seed” questions are inserted into the elicitation flow to gauge and adjust for response quality. Their role is mainly quality control, not to dominate the substantive target question.
3.5 Scenario-based elicitation
Scenario-based elicitation asks experts to assign probabilities or distributions across alternative plausible futures or conditions. This is useful when the system can plausibly follow distinct trajectories and when experts can reason more effectively in terms of scenarios than in terms of a single parametric model.
3.6 Dealing with dependent experts
When experts share information or collaborate, their judgments may be correlated. Dependence can be addressed by eliciting jointly structured inputs, collecting evidence about shared sources, using methods that explicitly model correlation, or restructuring the process to minimize feedback and discussion during elicitation.
4 Bias Reduction and Quality Control
4.1 Common judgment biases in scientific estimation
Human judgments can deviate from normative probability due to cognitive biases. Key examples include:
4.1.1 Anchoring and adjustment issues
Anchoring occurs when initial values influence subsequent estimates. To mitigate it, elicitation protocols often avoid suggesting plausible numbers, randomize order in which quantities are asked, and use careful wording so that experts begin from their own internal beliefs rather than an external anchor.
4.1.2 Overconfidence and under-dispersion
Overconfidence can lead experts to produce distributions that are too narrow, underrepresenting uncertainty. Under-dispersion can be detected through calibration checks and coherence tests. Protocols may counteract this by prompting wider quantile ranges, asking for multiple percentiles, and using scoring rules that penalize miscalibration.
4.2 Training and practice rounds
Practice rounds familiarize experts with the elicitation format and help reduce misunderstandings about probability scales and quantile interpretation. Training can also improve compliance with the expected structure, especially for complex distribution elicitation tasks.
4.3 Checking coherence and logical consistency
Coherence checks verify that the elicited quantities satisfy basic logical constraints, such as monotonicity of quantiles (e.g., lower percentiles should not exceed higher ones) and consistency between derived measures. Coherence can be evaluated before aggregation to identify errors caused by misunderstanding or calculation mistakes.
4.4 Uncertainty decomposition
Where appropriate, uncertainty can be decomposed into components, such as variability inherent to the phenomenon versus uncertainty due to limited knowledge. Decomposition helps interpret whether narrow distributions reflect genuine confidence or whether experts are collapsing different sources into a single undifferentiated range.
5 Aggregating Expert Judgments
5.1 Opinion pooling basics (linear, logarithmic, and other pools)
Aggregation combines individual expert beliefs into a group distribution. In linear pooling, distributions are averaged in probability space, while logarithmic pooling (a geometric-average form) multiplies likelihood-like components and tends to emphasize agreement. Different pooling rules can lead to different tail behavior and should be chosen with awareness of the decision context.
5.2 Weighting experts (performance-based vs. equal weighting)
Experts may be weighted to reflect reliability. Equal weighting is simple and often used when performance information is limited. Performance-based weighting can incorporate calibration or historical accuracy on comparable tasks. Weight selection should be justified and ideally supported by evidence rather than subjective preference.
5.3 Handling outliers and disagreement
Disagreement is informative, but extreme outliers can reflect data entry errors, misunderstandings, or miscalibration. Methods for handling disagreement include robustness adjustments, trimmed pooling, or models that treat outlier forecasts as coming from a different error process. The aim is to prevent undue distortion while preserving genuine diversity of belief.
5.4 Combining different evidence sources
Expert judgments may be integrated with empirical data, model outputs, or mechanistic simulations. This can be done via Bayesian updating, hierarchical models, or structured frameworks that treat expert beliefs as priors or likelihood surrogates. The combination should clarify how each evidence source contributes to final uncertainty.
6 Uncertainty Representation and Propagation
6.1 Choosing probability models
After elicitation, beliefs must be represented in a form suitable for downstream analysis. The choice of representation affects interpretability, computational tractability, and how uncertainty in tails is handled.
6.1.1 Parametric vs. nonparametric representations
Parametric representations use distributions characterized by a small set of parameters (e.g., lognormal or beta distributions), which can be efficient but may fail to fit complex shapes. Nonparametric representations rely on quantile functions, piecewise definitions, or flexible mixture models, often capturing skewness or multimodality more faithfully at the cost of additional complexity.
6.2 Uncertainty propagation to downstream metrics
Propagation refers to carrying elicited uncertainty through the analysis pipeline to obtain uncertainty on derived quantities. Techniques may include Monte Carlo simulation, numerical integration, or analytic approximations when feasible. The propagation step should preserve the relationship between uncertain inputs and the final metrics.
6.3 Reporting credible intervals and tail risk
Results are commonly summarized with central credible intervals (or similar interval measures) and with attention to tail risk when decision-makers face consequences from extreme outcomes. Tail reporting requires careful definition of exceedance thresholds and an understanding of how the chosen probability representation influences extreme probabilities.
7 Validating and Updating Elicited Beliefs
7.1 Backtesting against realized outcomes
Validation can occur after outcomes are observed by comparing forecasts to what actually happened. Backtesting evaluates calibration and discrimination—whether high-probability predictions correspond to frequent outcomes and whether the forecasts meaningfully separate likely from unlikely events.
7.2 Comparing to empirical data when it becomes available
As empirical evidence arrives, elicited beliefs can be checked against new measurements. Discrepancies reveal whether initial assumptions were incomplete, whether the elicitation process missed relevant information, or whether underlying models require modification.
7.3 Iterative updating (Bayesian and non-Bayesian approaches)
Updating uses new evidence to refine beliefs. Bayesian approaches incorporate evidence through likelihood functions and priors informed by elicitation. Non-Bayesian approaches may adjust distributions via rule-based updating, constraints, or optimization methods that incorporate observational summaries without fully specifying likelihoods.
7.4 Learning from calibration feedback
Calibration feedback informs both future elicitation design and expert training. If certain elicitation formats systematically underperform, the protocol can be revised (e.g., changing quantile sets, clarifying definitions, or improving coherence checks).
8 Reporting Standards and Reproducibility
8.1 Documenting assumptions and elicitation protocols
Reproducible reporting requires documenting the target quantity, question wording, quantile levels, elicitation order, background material provided to experts, and any scoring or calibration steps. Explicit protocols help readers understand why the elicited distribution has a particular shape.
8.2 Transparency about expert selection and incentives
Reports should describe criteria for expert recruitment, how independence was handled, and what incentives or scoring systems were used. Transparency allows assessment of potential bias introduced by selection or motivation.
8.3 Presenting results for different audiences
Different stakeholders may need different presentations: technical users may require full distributions, model parameters, and propagation methods, while decision stakeholders may need interpretable summaries such as risk thresholds and interval estimates. Good reporting aligns level of detail with the audience’s needs without hiding methodological choices.
8.4 Reproducible workflows and audit trails
Reproducible workflows include version-controlled data, documented transformation steps from elicited inputs to aggregated outputs, and audit trails of scoring, pooling, and uncertainty propagation. Such practices support verification and replication.
9 Applications and Use Cases
9.1 Forecasting rare events and emerging phenomena
In settings where rare events lack sufficient historical data, experts may be the only source for probability estimates. Elicitation provides structured quantification of low-frequency outcomes and helps incorporate mechanistic reasoning about why those events might occur.
9.2 Risk assessment and scenario planning
Risk assessment uses probabilistic judgments to evaluate potential harms under uncertainty. Scenario planning benefits from elicitation when the future depends on contingent developments, enabling explicit probability assignments to alternative plausible trajectories.
9.3 Scientific modeling when data are scarce
Early research frequently faces limited observations. Elicitation can guide model parameter priors, validate plausible ranges, and identify which assumptions dominate uncertainty, thereby informing where future data collection could be most valuable.
9.4 Policy-neutral analytical applications (e.g., research planning)
Elicitation can support research planning by quantifying uncertainty about experimental outcomes, timelines, or effect sizes. Even when not aimed at governance decisions, explicit uncertainty representations improve planning by clarifying expected variability and the likelihood of different outcomes.
10 Limitations and Challenges
10.1 Sensitivity to question wording
Elicited distributions can change when questions are phrased differently, especially if the target variable is ambiguous. Sensitivity analysis and iterative refinement of definitions help reduce the risk that the protocol measures interpretation rather than belief.
10.2 Measuring and ensuring expert independence
Ensuring independence is difficult when experts share networks or common training. Dependence can bias aggregated uncertainty if not handled appropriately. Documenting information sources and structuring the process can mitigate, though not always eliminate, correlation.
10.3 Representativeness and coverage of the expert community
A selected group of experts may not represent the broader domain knowledge distribution. Coverage issues can lead to systematic bias, particularly in fast-moving areas where expertise is unevenly distributed.
10.4 Practical constraints (time, cost, expertise availability)
Elicitation protocols can be time-consuming, especially when requiring training, calibration, and extensive quantile collection. Resource constraints may force shorter protocols, which can reduce quality. Practical planning balances rigor with feasibility.
11 Practical Step-by-Step Workflow
11.1 Plan the target quantity and uncertainty goal
Define the decision-relevant target variable, specify how uncertainty will be used downstream, and choose the output format needed (intervals, full distribution, or quantiles). Identify which uncertainty components matter for the analysis and set goals for calibration quality.
11.2 Conduct elicitation sessions
Run structured sessions following the written protocol. Provide consistent background information, ensure experts understand the scale and definitions, and use prompts or guided steps to obtain probabilistic answers. Where possible, include practice rounds and seed questions.
11.3 Validate, score, and calibrate judgments
Check coherence constraints, evaluate calibration using seed or validation questions, and apply scoring rules or adjustments consistent with the protocol. Address misunderstandings promptly through clarifying questions or retraining, when feasible.
11.4 Pool results and propagate uncertainty
Aggregate expert distributions using an appropriate pooling method and expert weighting strategy. Then propagate the resulting uncertainty through the model or analysis pipeline to compute uncertainty for downstream metrics, reporting both central intervals and relevant tail behavior.
11.5 Publish methods and update as new evidence arrives
Report question wording, expert selection, elicitation and scoring details, pooling choices, and propagation methods. After new evidence becomes available, update elicited beliefs and document changes so that the evolving inference remains traceable.