1 Setup and Notation

1.1 Triangular arrays of random variables

A common way to formulate extensions of the Central Limit Theorem beyond identically distributed variables is to use triangular arrays. One considers, for each integer \(n\), a finite sequence of random variables \[ \{X_{n,1},X_{n,2},\dots,X_{n,k_n}\}, \] where the row length \(k_n\) may depend on \(n\). The index \(n\) indicates the “stage” of approximation, while \(j\) identifies a term within that stage. The object of interest is typically the normalized sum of the entire \(n\)-th row.

1.2 Independence assumptions

The theorem is built for independent random variables within each row: for fixed \(n\), the variables \(X_{n,1},\dots,X_{n,k_n}\) are independent. Across different rows, no particular relationship is required for the statement as long as the limiting distribution is described correctly for each \(n\). Independence is crucial because it enables factorization of transforms and the control of cross-effects.

1.3 Normalization and partial sums

Let the row sum be \[ S_n=\sum_{j=1}^{k_n} X_{n,j}. \] In typical formulations, a normalization by the (row) standard deviation is applied to obtain a non-degenerate limit. One writes \( \sigma_n^2 \) for the total variance in row \(n\), and studies \[ \frac{S_n}{\sigma_n}. \] Sometimes the normalization is presented in terms of variance growth, or via a condition that implies \(\sigma_n^2 \to \sigma^2\).

1.4 Convergence in distribution and limiting normal law

The goal is asymptotic normality: there exists \(\sigma^2>0\) such that \[ \frac{S_n}{\sigma_n}\ \xrightarrow{d}\ \mathcal{N}(0,\sigma^2) \] (or equivalently a standard normal after scaling). Convergence in distribution means that expectations of bounded continuous test functions of the normalized sums converge to the corresponding expectations under the Gaussian law.

2 Statement of the Lindeberg–Feller Theorem

2.1 Mean and variance requirements

A typical version begins by requiring that the variables have (asymptotically) centered behavior and that the total variance stabilizes. Concretely, one assumes \[ \sum_{j=1}^{k_n}\mathbb{E}[X_{n,j}] \to 0 \] and \[ \sum_{j=1}^{k_n}\mathrm{Var}(X_{n,j}) \to \sigma^2, \] with \(\sigma^2>0\). Often the statement is simplified by taking each \(X_{n,j}\) to be centered, but the core logic only depends on the combined centering of the row sum.

2.2 Variance normalization and limiting variance

The normalization is commonly done using \[ \sigma_n^2=\sum_{j=1}^{k_n}\mathrm{Var}(X_{n,j}). \] The theorem then targets the distribution of \[ \frac{S_n}{\sigma_n}. \] When \(\sigma_n^2\to\sigma^2\), the limit law has the same variance parameter. The central idea is that the row sum behaves like a “Gaussian mixture of small independent increments,” provided no single increment overwhelms the total variance.

2.3 The Lindeberg condition

The defining requirement beyond mean and variance control is the Lindeberg condition, which limits the influence of rare, large values.

2.3.1 Truncation via ε-dependent thresholds

For every \(\varepsilon>0\), the Lindeberg condition requires that \[

\frac{1}{\sigma_n^2}\sum_{j=1}^{k_n}\mathbb{E}\!\left[X_{n,j}^2\mathbf{1}\{X_{n,j}>\varepsilon \sigma_n\}\right]\ \longrightarrow\ 0

\quad \text{as } n\to\infty. \] The indicator truncates contributions from “large” deviations—those whose magnitude exceeds a fraction \(\varepsilon\) of the row scale \(\sigma_n\). The condition says that, after normalizing by the total variance, the sum of the tail second-moments becomes negligible.

2.3.2 Tail contribution becoming negligible

Intuitively, the Lindeberg condition enforces a weak form of uniformity: even if variables are not identically distributed and may have different scales, the probability mass of excessively large outcomes must not contribute significantly to the quadratic energy that determines variance. This prevents a few extreme terms from dictating the limiting distribution.

2.4 Conclusion: asymptotic normality

Under the mean/variance assumptions and the Lindeberg condition, the normalized sums satisfy \[ \frac{S_n-\mathbb{E}[S_n]}{\sigma_n}\ \xrightarrow{d}\ \mathcal{N}(0,1). \] If one already has \(\mathbb{E}[S_n]=0\), then \(S_n/\sigma_n \Rightarrow \mathcal{N}(0,1)\). More generally, scaling yields the stated limiting variance \(\sigma^2\) when \(\sigma_n^2\to\sigma^2\).

3 Variance and Distributional Convergence

3.1 Role of variance convergence

Variance convergence anchors the scale of the limiting Gaussian law. If the total variance fails to stabilize (for example, tends to \(0\) or diverges without proper normalization), then the limiting distribution cannot be a non-degenerate normal. In practice, the theorem couples variance control with the Lindeberg condition so that second-moment behavior both exists and is meaningfully shared across the array.

3.2 Tightness and controlling fluctuations

To obtain a distributional limit, one must prevent the normalized sums from exhibiting erratic behavior at large scales. While the theorem’s statement is framed via characteristic functions and truncation arguments, the underlying effect is that Lindeberg’s requirement suppresses the chance that a small number of terms produce large normalized jumps. This yields tightness: mass does not escape to infinity in the normalized scale.

3.3 Connection to characteristic functions

A standard route to the theorem uses characteristic functions. For independent summands, \[ \varphi_{S_n/\sigma_n}(t)=\prod_{j=1}^{k_n}\mathbb{E}\left[e^{itX_{n,j}/\sigma_n}\right]. \] The exponential factors are expanded around \(t=0\), with the leading term governed by the second moments. The Lindeberg condition controls the remainder contributions from terms that would otherwise spoil the quadratic approximation.

3.4 Uniqueness of the Gaussian limit

Once the characteristic functions are shown to converge to those of a normal distribution, standard uniqueness results for characteristic functions imply that the limiting law is Gaussian. In this sense, the theorem not only proves existence of a normal limit under the stated hypotheses but also guarantees that the limit is uniquely determined by the limiting variance and the centering.

4.1 Lindeberg condition vs. Lyapunov condition

The Lyapunov condition is a stronger but simpler moment-based criterion. It typically assumes that for some \(\delta>0\), \[

\frac{1}{\sigma_n^{2+\delta}}\sum_{j=1}^{k_n}\mathbb{E}\left[X_{n,j}^{2+\delta}\right]\to 0.

\] Lyapunov implies Lindeberg (under standard centering), since higher moments control tails more efficiently. The trade-off is conceptual clarity versus sharpness: Lyapunov can be easier to verify but may exclude arrays where only second-moment tail control is available.

4.2 Uniform integrability interpretations

The Lindeberg condition can be interpreted through the lens of uniform integrability. The truncation formulation resembles the requirement that the family of squared variables, when normalized by the total variance, does not concentrate its mass in extreme regions. In many treatments, one rewrites the condition to show convergence of truncated second moments and the vanishing of tail contributions, which aligns with uniform integrability behavior.

4.3 “Conditional” viewpoints for arrays

Although the theorem is usually stated unconditionally, one can view it as a statement about how “most” of the variance arises from “typical” events rather than rare extremes. The truncation step effectively conditions on the event \(X_{n,j}\le \varepsilon\sigma_n\) and uses the fact that the excluded part becomes negligible. This “conditional” viewpoint helps clarify why the theorem generalizes the classical CLT: it replaces identical distribution with a condition on how variance is apportioned across the array.

4.4 Feller’s condition and further refinements

Feller’s condition is another related refinement, often expressed as a necessary and/or sufficient type constraint on the maximal contribution of individual terms to the total variance. In heuristic terms, it requires that no single \(X_{n,j}\) term retains a macroscopic share of the quadratic energy in the limit. While the precise statement varies by formulation, such conditions aim at controlling the “largest summand” behavior, complementing the Lindeberg requirement.

5 Proof Strategies (High-Level)

5.1 Characteristic function expansion

A central proof idea is to approximate each factor \[ \mathbb{E}\left[e^{itX_{n,j}/\sigma_n}\right] \] by a second-order expression when \(X_{n,j}\) is not too large. For small arguments, the exponential can be expanded: \[ e^{ix}\approx 1+ix-\frac{x^2}{2}. \] Summing over \(j\) converts the collection of second moments into the target Gaussian exponent. The remaining difficulty is showing that the approximation error vanishes as \(n\to\infty\).

5.2 Truncation and remainder control

To control the approximation error, proofs typically truncate each \(X_{n,j}\) at the level \(\varepsilon\sigma_n\). One works with truncated variables and demonstrates:

1 Setup and Notation

2 Statement of the Lindeberg–Feller Theorem

This yields a controlled remainder uniformly enough to pass to the limit.

5.3 Use of independence in the factorization step

Independence is used explicitly when taking products of characteristic functions. After truncation, the logarithm of the product becomes a sum of logarithms, and each term can be handled individually. Independence ensures there are no cross-covariances or higher-order interaction terms that would complicate the scaling.

5.4 Bounding error terms

Finally, one bounds the remainder terms produced by truncation and Taylor expansion. These bounds usually rely on:

  • the smallness of truncated variables relative to \(\sigma_n\),
  • the convergence of total variance,
  • the Lindeberg tail control.

Once all errors are shown to vanish, the limiting characteristic function is identified as that of a normal random variable.

6 Examples and Applications

6.1 Sums of heterogeneous independent terms

Consider a row where each \(X_{n,j}\) is an independent increment with a different distribution but with variances summing to a stable total. The theorem applies when each increment’s tail impact is negligible relative to the row variance. Such heterogeneity appears naturally in models where different components have different scales, such as varying noise levels across sensors or time steps.

6.2 Studentization and variance estimation (conceptual)

A conceptual application involves replacing the unknown normalization \(\sigma_n\) with an estimate derived from the data. The Lindeberg–Feller theorem provides the base asymptotic normality for properly normalized sums, and under additional assumptions (often ensuring the estimator converges appropriately) one can derive a normal limit for the studentized statistic. The theorem is thus an ingredient in asymptotic justification of many standardized procedures.

6.3 Randomized sampling and weighted sums

Weighted sums of independent terms arise when one selects components according to some sampling scheme and forms a sum with weights. If the resulting array can be expressed as independent summands with controlled second-moment behavior and Lindeberg tail suppression, then asymptotic normality follows. The theorem’s array framework is particularly suited to these situations because the number of terms and their distributions may vary with \(n\).

6.4 Classical CLT as a special case

The classical Central Limit Theorem is recovered when the array consists of i.i.d. centered variables \(X_{n,j}=Y_j\) with common variance and a growing number of summands. In that setting, the Lindeberg condition holds automatically under finite second moment, and the normalized sum converges to a standard normal. Thus, Lindeberg–Feller can be viewed as the exact mechanism that extends the CLT from identical distributions to arrays with potentially different laws.

7 Failure Modes and Intuition

7.1 When the Lindeberg condition breaks

If the Lindeberg condition fails, it means that for some \(\varepsilon>0\), the normalized sum of tail second moments does not vanish. In such cases, extreme values remain too influential after scaling, so the Gaussian approximation based on aggregated small increments becomes unreliable.

7.2 Dominance of a few large terms

A typical failure scenario is that a small number of terms in each row have variances that remain significant while occasionally producing very large outcomes. Then the limiting distribution can be non-Gaussian, reflecting jump-like behavior rather than the smoothing effect that yields a normal law. In effect, the “many small pieces” assumption implicit in CLT behavior is violated.

7.3 What convergence looks like without conditions

Without conditions controlling tails and variance distribution, different subsequences of sums can converge to different limits, or no non-degenerate limit may exist under the proposed normalization. The array might exhibit skewness, heavy-tailed stable behavior, or simply degenerate convergence if the scale is mis-specified. Lindeberg–Feller provides a clean criterion that rules out these pathologies for normal limits.

8 Connections to Other Results

8.1 Central Limit Theorem (classical) comparison

The classical CLT addresses i.i.d. sequences, while Lindeberg–Feller handles triangular arrays of independent, non-identically distributed variables. The Lindeberg condition is the replacement for identical distribution: it enforces a comparable balance of contributions across terms even when their distributions differ.

8.2 Martingale CLTs (conceptual contrast)

Martingale Central Limit Theorems treat dependent structures where conditional expectations and orthogonality replace independence. Conceptually, both frameworks provide asymptotic normality under a “no big jumps” regime and appropriate variance stabilization. The main difference is the dependence structure: Lindeberg–Feller uses independence and truncation, whereas martingale CLTs use conditional variance and martingale difference properties.

8.3 Stable convergence and generalizations

Generalizations often strengthen the notion of convergence and preserve more information about the limiting behavior. Lindeberg–Feller’s classical distributional convergence is a starting point; stronger modes like stable convergence can accommodate randomness in limiting variance or additional sigma-algebra information, while still relying on tail and variance control ideas reminiscent of Lindeberg’s.

Invariance principles connect functional limit behavior of processes to Gaussian limits. Although those results may require more machinery, the core theme overlaps with Lindeberg–Feller: show that appropriately normalized increments behave like Gaussian increments in the limit, typically by controlling the influence of extreme contributions and ensuring variance scaling is correct.