1. Definitions and Setup
1.1 Hypothesis testing framework and p-value construction
In a standard hypothesis testing problem, one considers a null hypothesis \(H_0\) and an alternative \(H_1\). A test statistic \(T\) is computed from the data, and its distribution under \(H_0\) is used to define the p-value. For an observed value \(t_{\text{obs}}\), a typical p-value is \[ p = \Pr_{H_0}\{T \ge T(t_{\text{obs}})\} \] for a right-tailed convention, or an analogous expression for left-tailed or two-sided conventions. Under regularity and continuity conditions, the p-value has a uniform distribution on \([0,1]\) when the null is true, which provides a baseline for calibration.
1.2 Noncentral alternatives and the role of noncentrality
When the true data-generating process deviates from the null in a systematic way, many classical models can be expressed using noncentrality parameters. For example, in Gaussian mean-shift settings, a nonzero signal-to-noise ratio yields noncentral normal representations for standardized statistics. In variance or quadratic-form settings, signal strength often induces noncentral chi-square distributions for sums of squares or likelihood-ratio-related statistics. The noncentrality parameter (often denoted \(\lambda\)) captures how far the true model is from the null in an “effective” sense; larger \(\lambda\) typically increases the probability of large (or extreme) values of \(T\), thereby increasing rejection likelihood.
1.3 Test statistics and tail conventions (one-sided vs two-sided)
| P-values depend on how evidence is summarized by \(T\) and on the tail convention. One-sided p-values typically correspond to a single-sided extremeness criterion relative to \(H_0\). Two-sided p-values treat departures in either direction as evidence, often through \(\Pr_{H_0}( | T | \ge | t_{\text{obs}} | )\) or equivalent formulations. Two-sided constructions may change both the exact distribution of p-values and their monotonic dependence on noncentrality, particularly when \(T\) has asymmetric behavior under the alternative. |
|---|
1.4 Assumptions: continuity vs discreteness and nuisance parameters
Two sources of deviation from “ideal” p-value behavior are (i) discreteness and (ii) nuisance parameters. If the test statistic is discrete, p-values defined via tail areas can become conservative (stochastically larger than uniform) because the null distribution has atoms. With nuisance parameters, the p-value may be conditioned on an estimate, profiled, or integrated out; each choice affects the distribution under alternatives. Under noncentrality, these effects can interact with the p-value definition, changing both the degree and shape of how often p-values fall below nominal thresholds.
2. Distribution of p-values Under Noncentrality
2.1 Stochastic behavior of p-values under alternatives
Under noncentral alternatives, the p-value distribution is generally not uniform. Instead, p-values tend to be stochastically smaller than under the null because the test statistic shifts toward the rejection region. “Stochastically smaller” can be expressed through inequalities on the survival function: for many tests and alternatives, \(\Pr(p \le \alpha)\) increases with noncentrality \(\lambda\). The precise form depends on the statistic, the p-value definition, and whether the p-value is one-sided or two-sided.
2.2 Relationship to power and rejection regions
The power function describes \(\Pr_{H_1}( \text{reject } H_0)\) as a function of the alternative parameter(s). For threshold-based testing, rejecting at level \(\alpha\) corresponds to events like \(p \le \alpha\) (when the p-value is calibrated to yield exact or conservative control under \(H_0\)). Consequently, \[ \Pr_{\lambda}(p \le \alpha) \] is closely tied to power under noncentrality. The p-value distribution under the alternative can therefore be read as a “continuum” of rejection probabilities across all possible \(\alpha\), not only at a single significance level.
2.3 Exact p-value distributions in canonical models
In canonical settings, the distribution of p-values can be derived explicitly using the known null and alternative distributions of \(T\). Exact formulas typically use noncentral cumulative distribution functions (CDFs) and their inverses.
2.3.1 Noncentral chi-square based tests
Many likelihood-based and quadratic-form tests involve statistics that, under the null, follow central chi-square distributions, while under alternatives follow noncentral chi-square distributions. When the p-value is defined as an upper-tail area of the central chi-square distribution, the distribution of the p-value under the noncentral chi-square alternative can be written by mapping each possible observed statistic value through the null tail function and then applying the noncentral distribution for \(T\). The resulting p-value distribution often becomes more concentrated near 0 as the noncentrality parameter increases.
2.3.2 Noncentral normal based tests
For standardized tests in location problems, the null distribution of a z-like statistic is normal, while under a mean-shift alternative it becomes a noncentral normal. When the p-value is defined using the null normal CDF tail(s), one can obtain the p-value distribution under the alternative through the monotone relationship between \(T\) and its null tail probability. In one-sided tests, this relationship is typically monotone, making it easier to track how p-values shrink with noncentrality.
2.4 Asymptotic p-value behavior with growing sample size
As sample size increases, many test statistics obey asymptotic normality or converge to limiting chi-square forms. Under contiguous or fixed alternatives, the effective noncentrality often grows with sample size (or stabilizes under local alternatives). As a result, p-values often become more extreme: under stronger signal regimes, \(\Pr(p \le \alpha)\) approaches 1 for any fixed \(\alpha\). Under local alternatives, the convergence can yield limiting p-value distributions governed by the limiting noncentrality in the asymptotic model.
3. Stochastic Ordering, Calibration, and Comparison
3.1 Stochastic dominance versus exact distribution
Two ways to characterize p-value behavior under alternatives are (i) exact distributional expressions and (ii) qualitative orderings. Exact distributions describe the full shape of the p-value CDF. Stochastic dominance, by contrast, provides inequalities such as \(\Pr(p \le \alpha)\) being larger than (or equal to) another test’s probability for all \(\alpha\). Dominance is especially useful when comparing procedures without requiring full closed-form p-value distributions.
3.2 Uniformity under the null versus deviation under alternatives
Under the null and continuity assumptions, p-values are uniform by construction. When alternatives induce noncentrality, deviations from uniformity can often be summarized by the CDF \(F_{\lambda}(u)=\Pr_{\lambda}(p \le u)\). Under typical noncentral alternatives, this CDF lies above the uniform CDF \(u\) (for most \(u\)) because smaller p-values occur more frequently. The magnitude of deviation reflects the test’s sensitivity and the strength of the noncentrality.
3.3 p-value monotonicity in noncentrality parameter
For many one-parameter noncentral families, p-value distributions are monotone in the noncentrality parameter: larger \(\lambda\) shifts the test statistic further into the rejection tail. For one-sided p-values built from monotone tail functions, this can imply that \(\Pr_{\lambda_2}(p \le \alpha)\) is at least as large as \(\Pr_{\lambda_1}(p \le \alpha)\) whenever \(\lambda_2>\lambda_1\). Two-sided p-values and discrete statistics can complicate monotonicity, potentially introducing non-smooth changes as the mapping between observed \(T\) and tail probability crosses discontinuities.
3.4 Comparing tests: how distributional differences affect thresholds
Different tests can have similar nominal levels under \(H_0\) yet differ materially under alternatives. Comparisons based solely on asymptotic power at a fixed \(\alpha\) may miss differences in the full p-value distribution. A test might yield p-values that are stochastically smaller than another across a range of thresholds, implying uniformly greater \(\Pr(p \le \alpha)\) for many \(\alpha\). Conversely, two tests can cross: one may be more effective at smaller \(\alpha\) while the other dominates around moderate thresholds. Distributional comparison therefore matters when interpreting p-values beyond a single decision rule.
4. Power Curves and p-value Thresholds
4.1 Mapping from p-value thresholds to critical regions
When p-values are derived from tail areas of a statistic \(T\), the event \(\{p \le \alpha\}\) corresponds to a critical region defined in terms of \(T\). This mapping can be explicitly characterized when the p-value construction is based on a monotone tail probability under \(H_0\). For continuous statistics, the critical value \(t_{\alpha}\) is typically determined by the null quantile, and the same region can be translated into a p-value cutoff. Under noncentrality, analyzing \(\Pr_{\lambda}(T \in \text{critical region})\) directly yields \(\Pr_{\lambda}(p \le \alpha)\).
4.2 Power as a function of noncentrality
Power can be expressed as a function of the noncentrality parameter, often written as \(\beta(\lambda)=\Pr_{\lambda}(\text{reject})\). In canonical models, the relationship can be computed via noncentral CDFs of the test statistic. As \(\lambda\) increases, power generally increases, though the rate of increase depends on the test statistic’s sensitivity to the alternative and the degrees of freedom involved in the chi-square or quadratic-form structure.
4.3 Interpreting “how small p-values get” under alternatives
A common practical question is not only whether one rejects at a chosen \(\alpha\), but how rapidly p-values concentrate near zero as signal strengthens. The p-value CDF \(F_{\lambda}(u)\) provides this information: steeper growth near \(u=0\) corresponds to more frequent very small p-values. This perspective complements power curves by revealing the distributional “shape” across the entire range of p-values, which can differ substantially between procedures even if they share similar power at one threshold.
4.4 Practical guidance for choosing significance levels under noncentrality
Because p-values are calibrated under \(H_0\), significance thresholds are fixed by the desired error control notion under the null. However, when the alternative is effectively active, practitioners often choose thresholds by anticipating how often p-values will be extremely small. This can guide the selection of \(\alpha\) in exploratory contexts or the choice of how stringent to be when comparing competing models. Care is needed: the desire for smaller p-values under an alternative should not be conflated with a validity statement about posterior evidence, and different tail conventions (one- vs two-sided) can materially change the observed p-value scale.
5. Special Cases and Common Test Statistics
5.1 Likelihood ratio and Wald-type test statistics
Likelihood ratio and Wald-type statistics are frequently connected to chi-square limits under the null, and noncentral chi-square representations under alternatives. The noncentrality parameter typically depends on the true parameter shift and the information matrix. As a result, both the distribution of the statistic and its induced p-value distribution can be computed using noncentral chi-square CDFs. In finite samples, exact distributions may require more careful treatment, but the noncentral framework often provides a strong approximation.
5.2 Score/LM tests and their noncentral representations
Score (LM) tests can also admit noncentral representations under alternatives, though the noncentrality parameter and the form of the limiting distribution can differ from those of likelihood ratio or Wald tests. Since LM statistics can be especially sensitive to certain model directions, their p-value distribution under noncentrality may concentrate differently. In practice, the choice of test affects not only power at a fixed significance level but also the distributional dynamics across p-value thresholds.
5.3 Permutation or bootstrap p-values under noncentral alternatives (conceptual behavior)
Permutation and bootstrap procedures typically aim to mimic the null distribution while allowing flexibility in dependence structures. Under noncentral alternatives, the observed test statistic is shifted, so the resulting p-values often become smaller, reflecting increased alignment of the observed statistic with the null-rejecting tail. The conceptual behavior remains similar to noncentral parametric models: p-values become more frequent near zero, though exact stochastic ordering may depend on how well the resampling scheme reproduces the null and how it handles nuisance parameters.
5.4 Discrete outcomes and conservative/anti-conservative effects
When test statistics are discrete (e.g., based on counts), p-values can exhibit steps rather than a smooth continuum. Under such discreteness, p-values defined by tail probabilities may be conservative under \(H_0\), meaning their distribution is stochastically larger than uniform. Under alternatives, the same discretization can lead to nonuniform distortions: sometimes p-values decrease in large jumps, producing anti-conservative-like behavior at certain thresholds, or otherwise limiting how small p-values can become given the coarse resolution of the statistic.
6. Simulation and Computation
6.1 Monte Carlo estimation of p-value distributions
A common approach to studying p-value distributions under noncentrality is simulation. One specifies a noncentral alternative model, generates many synthetic datasets, computes a p-value for each dataset using the null calibration, and then estimates the empirical CDF of the p-values. This yields direct estimates of quantities like \(\Pr(p \le \alpha)\) for a grid of \(\alpha\), and it provides an empirical view of stochastic dominance or crossing behavior between tests.
6.2 Numerical evaluation via noncentral CDFs
When canonical distributions apply and the test statistic has a known noncentral form, p-value distribution functions can be computed more directly. Numerically evaluating noncentral chi-square or noncentral normal CDFs (and associated tail probabilities under the null) can produce accurate estimates without resampling. This route is often faster and more reproducible, but it relies on correct modeling of the noncentrality and the appropriateness of the assumed null and alternative families.
6.3 Verifying stochastic ordering empirically
After obtaining simulated p-values for two or more procedures, empirical checks can test whether one test yields stochastically smaller p-values than another. This can be assessed by comparing empirical CDFs across thresholds, and by looking for systematic separation rather than only single-threshold effects. Because of Monte Carlo variability and discretization, it is common to use confidence bands or repeated simulation to distinguish real ordering from sampling noise.
6.4 Reporting simulation settings and reproducibility
Simulation studies should report the alternative parameterization (including the noncentrality value(s)), sample size, degrees of freedom, the exact p-value definition (one- vs two-sided, tail conventions), and how nuisance parameters are handled. Reporting the number of Monte Carlo replications, random seeds, and convergence diagnostics supports reproducibility. For bootstrap or permutation methods, additional details about resampling counts and whether p-values use mid-p adjustments or continuity corrections can be essential.
7. Interpretation and Common Misconceptions
7.1 “P-values under the alternative” versus “calibrating under the null”
A p-value is defined through the null distribution. Therefore, its interpretation is tied to how it behaves under \(H_0\), not under the alternative. Studying the distribution of p-values under noncentrality is useful for understanding power and for anticipating outcomes when the alternative is true, but it does not redefine what the p-value means as an evidence measure.
7.2 Misreading p-value distributions as posterior probabilities
The distribution of p-values under an alternative may tempt one to interpret small p-values as posterior odds in favor of \(H_1\). This is not generally valid. P-values are not posterior probabilities; without a specified prior model and a likelihood-based Bayesian framework, p-values do not directly encode \(\Pr(H_1\mid \text{data})\). Noncentrality-based analysis can explain why p-values are small when signals exist, but it does not convert frequentist output into Bayesian probabilities.
7.3 Handling two-sided p-values and sign conventions
Two-sided p-values aggregate evidence from both tails and can behave differently than one-sided p-values under noncentrality. In asymmetric situations or when the direction of departure matters for the test statistic, sign conventions influence how the tail event is defined and hence how p-values map to critical regions. Correct interpretation therefore requires specifying the tail construction and how it corresponds to the alternative direction.
7.4 When noncentral modeling is a good approximation
Noncentral frameworks are often approximations to finite-sample distributions. They are most credible when test statistics admit well-known asymptotic limits and when the noncentrality parameter reasonably captures the signal strength relative to noise and information. If the model is highly misspecified, the noncentrality parameter might be effectively mismatched, leading to inaccurate predictions of the p-value distribution. In such cases, simulation-based calibration using plausible alternatives may be more reliable.
8. Extensions
8.1 Composite alternatives and averaging over noncentrality
Real alternatives may form a composite set rather than a single noncentrality value. One can study “average” p-value behavior by integrating over plausible values of \(\lambda\) under a chosen weighting scheme (e.g., an assumed distribution on parameters). The resulting mixture distribution can be broader than any single-\(\lambda\) distribution, sometimes reducing the apparent concentration near zero if weight is placed on weak-signal regions.
8.2 Nuisance parameters and conditional/non-conditional p-values
With nuisance parameters, p-values may be constructed conditionally (conditioning on estimates or pivot quantities) or non-conditionally (integrating or profiling over nuisance structure). Under noncentral alternatives, nuisance handling can change the effective shift in the statistic and hence the p-value distribution. Conditional p-values may yield different stochastic behavior from marginal ones even when both are calibrated under the null, emphasizing that “under the alternative” behavior is method- and construction-dependent.
8.3 Multiple testing: behavior of p-values from correlated noncentral laws
In multiple testing, p-values from different hypotheses can be dependent, and each p-value may follow a distribution induced by its own noncentrality and by the dependence structure. The overall pattern of small p-values influences error rates and selection behavior. Under correlated alternatives, the joint distribution determines whether procedures that rely on independence approximations remain accurate. Noncentral analysis clarifies the mechanism by which dependence and signal strength jointly shape the observed p-value landscape.
8.4 Robustness to model misspecification (noncentrality mismatch)
If the assumed noncentral model uses a noncentrality parameter that differs from the true effective signal, predicted p-value distributions can be biased. Such mismatch may lead to under- or overestimation of how often p-values fall below thresholds. Robustness can be assessed by sensitivity analysis over a range of noncentrality values, by comparing with simulation under more faithful data-generating processes, or by using resampling methods that reduce reliance on precise parametric noncentral structure.