1 Definition and Intuition

1.1 Model misspecification and why “true” may not exist in-model

In many statistical applications, the analyst specifies a parametric model intended to approximate how the data are generated. Under misspecification, the actual data-generating mechanism lies outside the assumed family. In that situation there is no parameter value that perfectly reproduces the observed distribution, so it is misleading to think in terms of “the” true parameter inside the model.

1.2 Formal definition via optimization of expected loss

A pseudo-true parameter is defined as the parameter value within the assumed family that minimizes an objective determined at the population level. The objective is typically an expected loss computed with respect to the true (but unknown) data-generating distribution. The minimizing parameter is called “pseudo-true” because it is best with respect to the chosen criterion, not because it equals a genuine truth that lies within the model.

1.3 Connection to population objectives and limiting parameters

Because the pseudo-true value is defined through expectations under the true distribution, it can be viewed as a population target that governs what the estimation procedure converges to asymptotically. As sample size grows, random sample criteria often converge to their population counterparts, so optimizers from the sample problem typically approach the pseudo-true minimizer.

1.4 Relation to “best approximation” concepts

The term reflects an approximation perspective: the assumed model class plays the role of a set of candidate approximations to reality. Choosing the pseudo-true parameter amounts to selecting the closest element of that set according to a rule induced by the loss or divergence.

2 Common Mathematical Characterizations

2.1 Expected log-likelihood maximization

A central characterization arises when the analyst uses maximum likelihood estimation (MLE). If the model is misspecified, maximizing the expected log-likelihood yields a population optimum that becomes the pseudo-true value.

2.1.1 Kullback–Leibler divergence viewpoint

Under standard regularity conditions, expected log-likelihood maximization corresponds to minimizing a Kullback–Leibler (KL) divergence between the true data-generating distribution and the model distribution. Since KL divergence is nonnegative and equals zero only when the distributions match, this provides a natural information-theoretic sense in which the pseudo-true parameter is the “closest” model in that divergence.

2.1.1.1 Minimization of KL divergence and uniqueness considerations

The KL-minimizing parameter may be unique or may not be, depending on the model class and the shape of the objective function. When multiple parameter values produce the same closest distributional behavior, estimators may converge to a set or exhibit nonstandard limiting distributions, especially near boundaries or under lack of identifiability.

2.2 Estimating-equation (moment) definitions

Many estimators are defined by solving estimating equations of the form “sample moments equal zero.” In misspecification, the population version generally does not vanish at the true parameter (because the true parameter is not in-model), but it may vanish at a different parameter value. The pseudo-true parameter can therefore be characterized as the solution that makes the population estimating equations hold as closely as possible, often interpreted via projected moments.

2.2.1 Population moments and projected parameters

Let the estimating equation be built from functions whose expectations are computed under the model. Under misspecification, one can define a target parameter as the one for which the model-implied moments best align with the true moments in the sense induced by the estimating scheme. This yields a notion of “projection” from the true distribution onto the class of distributions generated by the model.

2.3 Risk minimization and M-estimation perspective

In the broader M-estimation framework, a pseudo-true parameter is simply the minimizer of an expected criterion (risk) induced by the estimating function. This view unifies estimation by likelihood, least squares, and many other procedures: the estimator targets the population risk minimizer even when the assumed model family is wrong.

2.4 Identifiability and parameterization effects

Pseudo-true values depend on what parameterization and model class are used. Even if two parameterizations represent essentially the same set of distributions, the pseudo-true parameter in terms of coordinates may differ. Additionally, if the model is not identifiable, the pseudo-true target may be a region rather than a single point, affecting both convergence and interpretation.

3 Asymptotic Behavior Under Misspecification

3.1 Consistency to the pseudo-true value

Under misspecification, many estimators remain consistent, but the target is the pseudo-true parameter rather than the genuine parameter of the data-generating mechanism. Consistency typically follows from uniform convergence of the sample criterion to its population analogue and from the pseudo-true parameter being a well-separated minimizer.

3.2 Convergence rates and regularity conditions

The rate of convergence and the limiting distribution depend on smoothness, curvature of the objective, and stochastic equicontinuity. Standard asymptotic normality results often require conditions such as differentiability of the criterion, existence of moments, and nonsingularity of relevant matrices at the pseudo-true point.

3.3 The “sandwich” covariance principle

A hallmark of misspecified inference is the use of a covariance estimator often described as a “sandwich.” It typically combines an outer-product-of-gradients term with an inverse Hessian-like term. This reflects the fact that variability and curvature are governed by different expectations when the model is wrong, so a covariance formula derived under correct specification generally underestimates uncertainty.

3.4 Interpretation of standard errors when assumptions fail

Standard errors computed under correct-model assumptions may be too optimistic or otherwise misleading under misspecification. Inference that uses the sandwich principle aims to correct this by estimating the appropriate variability around the pseudo-true target, thereby producing uncertainty quantification that better reflects model mismatch.

3.5 Robustness limits and when pseudo-true inference matters

Pseudo-true inference is most meaningful when the misspecification is structural but the estimator still targets a stable minimizer and asymptotic approximations remain accurate. If mismatch is severe (e.g., leads to infinite variance, lack of concentration, or multimodality), pseudo-true-based asymptotics can break down and diagnostics and alternative modeling strategies become necessary.

4 Examples and Worked Scenarios

4.1 Linear regression with incorrect error distribution

Consider linear regression where the mean function is correctly specified but the error distribution is not (for instance, heavy-tailed or heteroskedastic errors). Even if maximum likelihood assumes Gaussian errors, the pseudo-true parameter for the regression coefficients typically aligns with the best approximation of the conditional mean under the assumed likelihood. The variance of the estimator then depends on the true error structure, motivating robust covariance estimates.

4.2 Generalized linear models under wrong likelihood form

In generalized linear models, misspecification can occur when the assumed link function or variance relationship is incorrect. The pseudo-true parameter is defined by the population risk minimizer induced by the wrong likelihood. Estimators converge to this parameter, while the usual model-based covariance matrix may be incorrect, again pointing to sandwich-type uncertainty quantification.

4.3 Misspecified mixture or nonlinear models

For nonlinear models or mixture models, the pseudo-true target may reflect an approximation that balances incompatible distributional features. Optimization landscapes can be complex, and the pseudo-true set can include multiple local minimizers. As a result, different estimation runs or initialization strategies may converge to different pseudo-true values, especially near identifiability boundaries.

4.4 Comparing pseudo-true values across alternative model classes

When comparing two model classes, each class yields its own pseudo-true parameter according to the criterion it optimizes. Differences between these pseudo-true parameters can be interpreted as a model-induced approximation gap rather than as a failure of estimation alone. Such comparisons help separate “statistical” limitation (information about the parameter) from “modeling” limitation (how close the class can get to the true distribution).

4.5 Simulation guidance for observing pseudo-true limits

Simulation can reveal pseudo-true behavior by generating data from a known mechanism outside the fitted model family, then repeatedly fitting the misspecified model. As sample size increases, the fitted parameters typically stabilize near the pseudo-true value. Observing whether estimates concentrate around a stable point, a set, or a drifting value helps diagnose whether assumptions for asymptotic theory are plausible.

5.1 True parameter vs pseudo-true parameter

The true parameter refers to a parameter value that exactly characterizes the data-generating distribution within the assumed model. The pseudo-true parameter replaces this role when exact representation is impossible, providing the closest in-model parameter under a chosen criterion.

5.2 Quasi-true parameter terminology

In some literature, “quasi-true” is used similarly to “pseudo-true,” especially in contexts emphasizing misspecified likelihood or moment conditions. Terminology can vary across communities, but the core idea remains: the estimation target is determined by a population optimization or moment-projection under misspecification.

5.3 Projection in function spaces (general approximation viewpoint)

Pseudo-true parameters can be linked to projection ideas in abstract spaces of distributions or functions. The model class is treated as a constrained subset, and the pseudo-true element arises from projecting the true object onto that subset using an induced notion of distance or discrepancy.

5.4 Fisher consistency vs pseudo-consistency

Fisher consistency is a property that, under correct specification, the functional targeted by an estimator returns the true parameter. Under misspecification, the analogous notion is pseudo-consistency: the estimator converges to the pseudo-true parameter defined by the population criterion rather than to any nonexistent “true-in-model” value.

5.5 Bayesian analogues (misspecified priors and posterior concentration)

Bayesian methods under misspecification often concentrate around parameters or sets that maximize expected log-likelihood or minimize an information divergence, but the limiting behavior depends on both the likelihood misspecification and the prior. Posterior credible intervals may require adjustments because nominal Bayesian uncertainty calibrated under correct specification may not match frequentist coverage under the real data-generating process.

6 Practical Implications for Inference

6.1 What parameter is targeted by maximum likelihood in misspecification

When using maximum likelihood with a misspecified model, the target is the KL-minimizing parameter (or more generally, the expected log-likelihood maximizer) within the assumed family. Thus, interpretations of fitted coefficients should be framed as best approximations under the modeling assumptions embedded in the likelihood.

6.2 Model diagnostics linked to pseudo-true gaps

Diagnostics can be interpreted as tools for detecting large pseudo-true gaps—situations where the best achievable fit within the model class is still far from the data-generating distribution. While diagnostics cannot identify the pseudo-true value directly (since the true mechanism is unknown), they can reveal systematic patterns inconsistent with the model’s approximation capacity.

6.3 Robust estimation and adjustment strategies

A common response is to use robust covariance estimates that reflect misspecification, such as sandwich estimators. Depending on the context, analysts may also employ alternative loss functions, robust M-estimators, or model re-specification to reduce the discrepancy driving the pseudo-true target.

6.4 Interpretation of hypothesis tests under misspecification

Hypothesis tests that rely on standard likelihood-based reference distributions may be distorted when the model is wrong. Pseudo-true asymptotics clarify that test statistics should be centered and scaled relative to the pseudo-true target, and that variance and curvature estimates should be adjusted accordingly to avoid misleading p-values.

6.5 Goodness-of-fit metrics and their relation to expected loss

Goodness-of-fit criteria often approximate expected loss or predictive discrepancy measures. Under misspecification, improvements in such metrics can be understood as reducing the population objective defining the pseudo-true parameter, rather than as recovering a nonexistent exact truth within the model.

7 Estimation and Computation

7.1 Solving for the pseudo-true parameter numerically

Because pseudo-true parameters are defined by population expectations that are unknown, in practice one approximates them using sample analogues. With large samples, optimizing the empirical criterion provides an estimator that converges to the pseudo-true parameter under regularity assumptions.

7.2 Profile likelihood and optimization strategies

For models with nuisance parameters, profile likelihood techniques reduce dimensionality by optimizing over nuisance parameters first. Under misspecification, profile likelihood still converges toward a population-profile objective, though the curvature and resulting confidence procedures must incorporate misspecification-aware variability.

7.3 Implementation with estimating equations

When the estimator is defined by estimating equations, computing the pseudo-true target corresponds to solving the sample version of the population moment conditions. Numerical solvers such as Newton-type methods require careful handling of step sizes and convergence tolerances, particularly when the objective has flat directions.

7.4 Bootstrap and resampling considerations under misspecification

Bootstrap methods can provide empirical standard errors or interval estimates, but validity may depend on how the resampling scheme reflects the misspecification. Some bootstrap variants are designed to approximate the sandwich variance behavior, while others may inherit the wrong-model assumptions and thus require correction or alternative resampling strategies.

7.5 Numerical stability and verification steps

Misspecification can lead to poorly conditioned optimization problems, especially in mixture or highly nonlinear models. Practical verification includes checking sensitivity to initialization, monitoring convergence diagnostics, evaluating gradients and Hessians at solutions, and performing out-of-sample predictive checks aligned with the pseudo-true objective.

8 Theoretical Refinements

8.1 Local misspecification and perturbation expansions

A refined approach considers misspecification as small perturbations of a correctly specified model. Perturbation expansions quantify how the pseudo-true parameter deviates from the true parameter and how standard asymptotic results deform as the degree of mismatch increases.

8.2 Semiparametric efficiency under misspecification

In semiparametric models, misspecification may occur either in the parametric component or in the remaining infinite-dimensional part. The efficient estimator under correct specification may no longer minimize the relevant asymptotic variance. Theory then identifies which modified efficiency notions apply when the target is pseudo-true rather than true-in-model.

8.3 Misspecification in dependent data settings (high level)

When observations are dependent (time series, spatial data, network data), misspecification interacts with dependence structure. Asymptotic results must account for altered variance due to both model mismatch and dependence, often requiring mixing or dependence-adjusted empirical process tools.

8.4 Uniform convergence and empirical process conditions

A deeper justification of pseudo-true convergence relies on uniform convergence of empirical criteria and control of stochastic fluctuations. Empirical process conditions ensure that sample objectives track population objectives well across parameter values, supporting consistency and asymptotic normality.

8.5 Multiple optima and boundary cases

When the pseudo-true objective has multiple minimizers or when minimizers lie on the boundary of the parameter space, the asymptotic distribution of estimators can deviate from the usual normal form. Careful treatment may involve set-valued limits, mixture distributions, or different scaling regimes.

9 Summary and Key Takeaways

9.1 Core definition and how it is determined

The pseudo-true parameter is the best approximating parameter inside a misspecified model family, determined by minimizing an expected loss or an information-theoretic divergence.

9.2 What estimators converge to

Under misspecification, many estimators still converge, but typically to the pseudo-true parameter rather than to any true parameter that would require correct model specification.

9.3 How uncertainty quantification is adjusted

Because curvature and variability can mismatch under incorrect assumptions, uncertainty quantification often needs misspecification-aware covariance formulas, frequently of the sandwich type, to better reflect true variability around the pseudo-true target.

9.4 When pseudo-true inference is reliable and informative

Pseudo-true-based inference is most reliable when optimization targets a well-separated pseudo-true minimizer and when regularity conditions for asymptotic theory hold. When mismatch is extreme or optimization is unstable, diagnostics and alternative modeling approaches become essential.