1 Definition and Intuition
1.1 Scoring rules and probabilistic forecasts
A scoring rule assigns a numerical score to a forecaster’s reported probability distribution and the eventual realized outcome. In probabilistic forecasting, the report typically takes the form of a distribution over possible outcomes, and the scoring rule evaluates how well that distribution anticipated the outcome that actually occurred. The forecast is assessed through the expected score computed under the forecaster’s believed data-generating distribution.
1.2 Properness vs. strict properness
A scoring rule is proper if truthful reporting maximizes expected score among all reports. Strict properness strengthens this: the truthful report is not only optimal, but is the *only* report (up to relevant equivalences) that achieves the maximum expected score. Intuitively, strict properness eliminates ties that could allow a forecaster to change their report without reducing expected performance.
1.3 Unique maximization of expected score
Let \(P\) denote the forecaster’s true belief (the distribution governing the next outcome) and let \(Q\) be a reported distribution. Strict properness means that the expected score \( \mathbb{E}_{X\sim P}[S(Q,X)] \) is maximized uniquely at \(Q=P\). “Uniquely” is important: if some other \(Q'\neq P\) could achieve the same expectation, then the incentive to report truthfully may be weaker, especially in settings where forecasters face additional strategic or practical constraints.
1.4 Incentive compatibility in forecasting
In mechanism design language, strict properness corresponds to incentive compatibility with a unique best response under the forecaster’s true beliefs. If a scoring rule is strictly proper, then a rational forecaster expecting future outcomes drawn from \(P\) maximizes their long-run average score by reporting \(P\). This provides a formal guarantee that evaluation protocols based on the rule reward honest probabilistic beliefs.
2 Mathematical Formulation
2.1 Expected score under a true distribution
Consider an outcome space \(\Omega\) and a class of distributions \(\mathcal{P}\) over \(\Omega\). A scoring rule \(S(Q,\omega)\) assigns a real number when outcome \(\omega\) occurs and the forecaster reports \(Q\). If the forecaster’s true belief is \(P\), the relevant quantity is the expected score \[ \mathbb{E}_{\omega\sim P}[S(Q,\omega)]. \] The scoring rule’s optimality properties are determined by how this expectation varies with the report \(Q\).
2.2 Reported distribution and optimization objective
Strict properness can be written as an optimization statement: for each true \(P\in\mathcal{P}\), \[ \mathbb{E}_{\omega\sim P}[S(P,\omega)] \;>\; \mathbb{E}_{\omega\sim P}[S(Q,\omega)] \quad \text{for all } Q\in\mathcal{P},\, Q\neq P, \] where strictness means a strict inequality for every distinct report. In many formulations, the strictness conclusion may hold for all \(Q\) that are distinguishable in the region where the forecaster’s score is sensitive (for example, where likelihoods are nonzero).
2.3 Inequalities that define strict properness
A closely related way to state strict properness is through an inequality involving a gap between the best expected score and any misreport. Define \[ D(P,Q) := \mathbb{E}_{\omega\sim P}[S(P,\omega) - S(Q,\omega)]. \] Then strict properness corresponds to \(D(P,Q)>0\) for all \(Q\neq P\) (under appropriate support conditions). This gap interpretation is useful in theoretical work because it quantifies the expected loss from deviating from the truth.
2.4 Relation to Bayes risk and minimization
In statistical decision theory, reporting and scoring can be linked to Bayes risk. Many scoring rules correspond to minimizing an expected loss \(L(Q,\omega)\) where \(S=-L\) up to sign conventions. Strict properness then becomes a statement that the Bayes action is unique: the report that minimizes expected loss under true \(P\) is the truthful distribution \(P\). This yields a direct bridge from probabilistic forecasting to optimization principles.
3 Characterizations
3.1 Convex analysis perspective
Strict properness admits a geometric interpretation in terms of convexity. A common characterization is that proper scoring rules correspond to convex functionals on the space of distributions, with strict properness corresponding to strict convexity (again, modulo subtleties like equivalence classes or support restrictions). Under this view, expected scores behave like values of a supporting hyperplane or tangent plane, and strictness reflects curvature that prevents “flat” directions where multiple reports tie.
3.2 Connection to Bregman divergences
A major family of results connects proper scoring rules to Bregman divergences. In such characterizations, the expected score difference \(D(P,Q)\) can be expressed as a Bregman divergence generated by a convex potential function. Strict properness corresponds to the associated divergence being strictly positive when \(Q\neq P\). This framework is valuable because it provides both intuition (divergence as “distance-like” expected loss) and algebraic tools for verifying strictness.
3.3 Bayes acts and uniqueness
Decision-theoretically, a scoring rule induces a set of optimal reports for each true belief \(P\). Strict properness is equivalent to the set of Bayes-optimal actions containing exactly one element (again, up to conditions where reports are behaviorally indistinguishable). This “uniqueness of Bayes act” perspective emphasizes that strict properness is a statement about the structure of optimal decision rules, not merely about local optimality.
3.4 Differentiable forms and gradient conditions
When scoring rules admit differentiability with respect to the reported distribution (or a parameterization of it), strict properness can be checked using derivative conditions. In differentiable settings, optimality corresponds to vanishing gradients at \(Q=P\), while strictness corresponds to a positive-definite curvature condition (often expressible through a Hessian being positive definite on the relevant domain). These analytic criteria are frequently used for parametric models, where one searches for unique maxima in expected score.
4 Common Examples of Strictly Proper Scoring Rules
4.1 Logarithmic score
The logarithmic score evaluates a report \(Q\) by rewarding the log-probability assigned to the realized outcome: \[ S(Q,\omega)=\log q(\omega), \] with appropriate conventions when probabilities are zero. Under standard support assumptions, the expected log score is uniquely maximized by reporting the true distribution. The strictness arises because the expected log score difference corresponds to a (nonnegative) divergence, and it vanishes only when the two distributions coincide.
4.2 Quadratic (Brier) score
The Brier score uses squared error between predicted probabilities and the realized outcome indicator. For a finite outcome space, if the report is \(Q\) and the realized outcome is \(\omega\), \[ S(Q,\omega)=2q(\omega)-\sum_{i} q(i)^2 \] (up to additive constants and scaling conventions). The associated expected score gap becomes a quadratic form that is minimized uniquely at the truth under full-support or standard domain assumptions. This rule is widely used because it is simple and yields interpretable penalties for probability misallocation.
4.3 Spherical score (and related forms)
The spherical score generalizes quadratic-like behavior while retaining properness. In a typical formulation on a finite space with probabilities \(q(i)\), \[
| S(Q,\omega)=\frac{q(\omega)}{\|q\|} |
|---|
\]
| where \(\|q\|\) denotes an \(\ell_2\)-norm. For appropriate domains (e.g., restricting to strictly positive probability vectors), expected score is uniquely maximized by the true distribution. Compared with the Brier score, spherical scoring can weight discrepancies differently across regions of the simplex. |
|---|
4.4 Rank-weighted and transformed scores
Beyond the classic examples, there are rank-weighted or transformed scoring rules designed to emphasize particular parts of the distribution, such as tail events or relative ordering. Many such constructions remain strictly proper when the transformation preserves the convex-analytic structure underlying properness. Verifying strictness often requires checking that the induced divergence remains strictly positive for distinct reports.
5 Discrete vs. Continuous Settings
5.1 Strict properness on finite outcome spaces
On finite outcome spaces, strict properness is typically easier to verify because the space of distributions is finite-dimensional (simplexes). Under standard assumptions (e.g., allowing all strictly positive distributions when needed), many common scoring rules—logarithmic, Brier, and spherical—are strictly proper. Uniqueness can be characterized by algebraic positivity of the associated expected loss gap.
5.2 Conditions for continuous distributions
For continuous outcomes, distributions may be represented by densities relative to a dominating measure, and care is needed with events of probability zero. Logarithmic scoring commonly requires ensuring that the true density and the reported density are compatible (e.g., the report does not assign zero density where the true density is positive). Strict properness can hold, but only within a domain of reports where the expected score is well-defined and where distinctions between distributions are measurable through the score.
5.3 Measure-theoretic considerations
In measure-theoretic formulations, strict properness is stated in terms of distributions on measurable spaces and often involves absolute continuity conditions. For log-type scores, integrability conditions can matter: if a candidate report places too little mass in regions critical under the true distribution, the expected score may be undefined or can diverge to \(-\infty\). Proper strictness statements therefore include conditions ensuring the objective is finite and the “distance-like” divergence is meaningful.
5.4 Handling constraints on reported distributions
Strict properness is sensitive to constraints. If forecasters are restricted to a smaller family of distributions—such as parametric models—then the unique maximizer may be the closest model in an appropriate divergence sense, not necessarily the true distribution itself. The rule can remain strictly proper *within the restricted class*, meaning it uniquely identifies the best report allowed by the constraints, rather than uniquely recovering the full true distribution.
6 Verification and Testing Strictness
6.1 Checking strict inequalities
A direct way to verify strict properness is to establish that the expected score difference is strictly positive for all distinct reports in the allowable domain. This often reduces to showing strict positivity of an associated divergence or quadratic form. The verification can be algebraic (as in finite-dimensional quadratic scores) or integral-based (as in log scores and measure-theoretic versions).
6.2 Identifiability and uniqueness issues
Even if a rule appears strictly proper, uniqueness can fail due to identifiability problems: distinct reports might induce identical score behavior over the events that the rule can “see.” For example, two density functions that agree almost everywhere (with respect to the relevant measures) may be treated as equivalent under the scoring rule. Strict properness is therefore commonly defined modulo such equivalences.
6.3 Sensitivity to modeling assumptions
Strictness may depend on assumptions about admissible reports, support, and finiteness of expectations. A scoring rule might be proper but not strictly proper if the scoring function is insensitive to certain degrees of freedom in the report. Conversely, adding domain restrictions can restore strictness by removing directions in which the score is flat.
6.4 Empirical approximation and numerical stability
In practice, strict properness is assessed with finite data and approximated expectations. Numerical estimation can blur strict inequalities, making it harder to distinguish near-equal expected scores. Care is needed to avoid artifacts from sampling variability, regularization, or constrained optimization when verifying that the chosen scoring rule genuinely encourages unique truthful reporting under the intended model class.
7 Extensions and Related Concepts
7.1 Proper scoring rules for composite events
Often, interest centers on composite events, such as functions of outcomes or derived quantities. Proper scoring rules can be extended so that the forecaster reports probabilities for these composite events while still receiving incentives aligned with truthful beliefs about the underlying distribution. Whether the extension is strictly proper depends on whether the mapping from underlying distributions to composite-event probabilities preserves identifiability.
7.2 Elicitation with restrictions e.g., parametric reports
When forecasters must report within a restricted family (for instance, a parametric form like a Gaussian with unknown mean and variance), strict properness becomes restricted strict properness. The scoring rule is strictly proper relative to the restricted class if the best expected score report within that class is unique and corresponds to the appropriate Bayes act under the belief. This is central in statistical practice, where full distributional flexibility is rarely feasible.
7.3 Multi-outcome generalizations
For settings with many outcomes, scoring rules generalize by acting on the full probability vector. Strict properness typically requires that the induced divergence or curvature remains strictly positive when two probability vectors differ. In high-dimensional cases, maintaining strictness can interact with numerical constraints and with whether the forecast space includes all relevant probability assignments.
7.4 Links to calibration and forecast evaluation
Strict properness is closely related to forecast evaluation concepts such as calibration. While calibration focuses on whether reported probabilities match observed frequencies, strict properness focuses on whether the scoring rule’s expectation is maximized uniquely by the true distribution. In many evaluation frameworks, strict properness helps ensure that average score improvements correspond to genuine improvement in probabilistic accuracy rather than merely reweighting outcomes without enhancing beliefs.
8 Practical Implications for Forecasting
8.1 Choosing scoring rules for model comparison
Scoring rules provide a basis for comparing forecasting models by evaluating predictive distributions on realized outcomes. Strict properness is desirable because it prevents models from improving scores through strategic distortion: higher expected scores correspond to better alignment with true underlying beliefs. This supports fair comparisons, especially when models are flexible and can otherwise exploit scoring artifacts.
8.2 Avoiding strategic misreporting
When forecasts are submitted by agents whose beliefs may differ from their reports, strict properness functions as an incentive mechanism. Under ideal assumptions, a rational agent gains no advantage from misreporting because any deviation lowers expected score. This reduces the need for ad hoc penalties or complex verification procedures aimed at preventing manipulation.
8.3 Interpretation of average scores
The average score over repeated trials is an empirical proxy for expected score. Under strict properness, improvements in expected score correspond to moving closer to the true distribution in the divergence sense induced by the scoring rule. However, interpreting differences between observed averages requires accounting for sampling variability, especially when outcome probabilities are extreme.
8.4 Trade-offs robustness vs strictness
In some applications, forecasters and evaluators seek robustness to model misspecification, heavy tails, or outliers. Strictly proper rules like the log score can be sensitive to probability mass assigned near zero, which may increase variance or produce large penalties. Less aggressive proper rules can offer stability at the cost of reduced curvature, sometimes affecting strictness depending on the domain. The choice of scoring rule therefore involves a balance between incentive strength and practical reliability.