1 Convexity and supporting hyperplanes

1.1 Proper convex functions and domains

Let \(X\) be a real vector space, typically equipped with enough structure to speak about continuous linear functionals (e.g., a locally convex topological vector space). A function \(f:X\to\mathbb{R}\cup\{+\infty\}\) is convex if \[ f(\lambda x+(1-\lambda)y)\le \lambda f(x)+(1-\lambda)f(y) \] for all \(x,y\in X\) and \(\lambda\in[0,1]\). It is proper if it is not identically \(+\infty\) and never takes the value \(-\infty\). The effective domain \(\mathrm{dom}\,f=\{x\in X:\ f(x)<+\infty\}\) is the region where the function is finite; many subdifferential statements depend on whether a point lies in, on, or outside this set.

1.2 Epigraph and geometric interpretation

The epigraph of \(f\) is \[ \mathrm{epi}\, f=\{(x,t)\in X\times\mathbb{R}:\ t\ge f(x)\}. \] Convexity of \(f\) is equivalent to convexity of its epigraph. In this geometry, supporting hyperplanes to \(\mathrm{epi}\,f\) at a point above the graph encode “best linear underestimates” of the function.

1.3 Supporting hyperplane characterization

A supporting hyperplane to the epigraph at a point \((\bar x,\bar t)\) is a hyperplane that touches \(\mathrm{epi}\,f\) there and lies entirely below it in the \((x,t)\)-direction. If such a hyperplane exists at \((\bar x,f(\bar x))\), it corresponds to a linear functional that does not exceed the graph of \(f\) at any point, and matches the function at \(\bar x\) in a directional sense. The set of all such slopes is precisely what the subdifferential formalizes.

1.4 Local versus global supporting slopes

For convex functions, the notion of support at \(\bar x\) is global even though it is defined locally via an inequality at all points \(x\in X\). However, in many cases it suffices to check the supporting inequality on a neighborhood of \(\bar x\), because convexity propagates local bounds into global ones. This distinction is important in theorems about existence and exactness of calculus rules.

2 Definition of the subdifferential

2.1 Subdifferential at a point

For a proper convex function \(f:X\to\mathbb{R}\cup\{+\infty\}\) and a point \(\bar x\in \mathrm{dom}\,f\), the subdifferential of \(f\) at \(\bar x\) is the set \[ \partial f(\bar x)=\{\ell\in X^*:\ f(x)\ge f(\bar x)+\ell(x-\bar x)\ \text{for all }x\in X\}, \] where \(X^*\) denotes the space of continuous linear functionals on \(X\).

2.1.1 Supporting inequality formulation

The defining inequality can be rewritten as \[ f(x)\ge \ell(x)+\bigl(f(\bar x)-\ell(\bar x)\bigr), \] showing that every \(\ell\in \partial f(\bar x)\) generates an affine function that lies below \(f\) and matches the value at \(x=\bar x\). Thus, \(\partial f(\bar x)\) collects all linear “supporting slopes” to the graph at \(\bar x\).

2.1.2 Set-valued mapping properties

The map \(x\mapsto \partial f(x)\) is set-valued: typically it is empty outside \(\mathrm{dom}\,f\) and may contain many elements on nondifferentiable points. It is nevertheless highly structured: for convex \(f\), subdifferentials are convex sets, and they interact well with the geometry of the epigraph and with monotone operator theory.

2.2 Subdifferential of the entire function

Sometimes one views \(\partial f\) as an operator that assigns to each \(x\in X\) its subdifferential \(\partial f(x)\). The graph of this operator is \[ \mathrm{gph}\,\partial f=\{(x,\ell)\in X\times X^*:\ \ell\in \partial f(x)\}. \] Many analytical properties (e.g., monotonicity) are expressed directly in terms of this graph.

2.3 Fenchel–Young notation and conventions

In Fenchel duality, one uses the pairing \(\langle \ell, x\rangle\) between \(X^*\) and \(X\) (or an analogous bilinear form depending on the setting). With this convention, the subgradient inequality is \[ f(x)\ge f(\bar x)+\langle \ell, x-\bar x\rangle. \] Conventions about extended values \(+\infty\) ensure inequalities remain meaningful even at points outside the effective domain.

2.4 Differentiable case as a special instance

If \(f\) is differentiable at \(\bar x\) (in the sense of a gradient compatible with the space), then the subdifferential collapses to a singleton: \[ \partial f(\bar x)=\{\nabla f(\bar x)\}. \] More generally, when \(f\) is subdifferentiable and admits a linearization at \(\bar x\), that linearization is exactly the unique supporting functional, so the subdifferential provides a consistent generalization of the derivative.

3 Basic properties

3.1 Nonemptiness conditions

Subdifferentials can be empty if the function has certain types of boundary behavior. A common theme is that interior-type points of the domain yield existence of supporting hyperplanes, while boundary points may fail to do so without additional assumptions.

3.1.1 Interior-point criteria

In finite-dimensional spaces, if \(f\) is convex and \(\bar x\) lies in the interior of \(\mathrm{dom}\,f\), then \(\partial f(\bar x)\) is nonempty under standard hypotheses. In more general topological vector space settings, analogous interior or regularity conditions (often involving relative interiors) guarantee the presence of at least one supporting functional.

3.2 Convexity and closedness aspects

For a fixed \(\bar x\), the set \(\partial f(\bar x)\) is convex: if \(\ell_1,\ell_2\in \partial f(\bar x)\), then any convex combination \(\lambda \ell_1+(1-\lambda)\ell_2\) also supports \(f\) at \(\bar x\). Depending on the ambient space and continuity assumptions, \(\partial f(\bar x)\) is also typically closed in the dual topology. These properties align with the hyperplane picture: supporting slopes form a convex family.

3.3 Monotonicity of the subdifferential

A central structural fact is that \(\partial f\) is a monotone operator: whenever \(\ell_1\in \partial f(x_1)\) and \(\ell_2\in \partial f(x_2)\), \[ \langle \ell_1-\ell_2,\, x_1-x_2\rangle \ge 0. \] This inequality follows from applying the subgradient inequality at both points and adding the resulting constraints. Monotonicity underpins stability and convergence results in variational analysis and optimization.

3.4 Scaling and translation rules

Subdifferentials respect simple transformations of \(f\). For example, if \(a>0\), then \[ \partial (a f)(x)=a\,\partial f(x). \] Adding an affine function \(\varphi(x)=\langle \ell_0,x\rangle + c\) shifts subgradients by \(\ell_0\): \[ \partial (f+\varphi)(x)=\partial f(x)+\{\ell_0\}. \] Vertical shifts \(f\mapsto f+c\) do not change \(\partial f(x)\), since the subgradient inequality depends on differences.

3.5 Examples: norms, indicators, and piecewise-linear functions

- Norms: For \(f(x)=\|x\|\) in a normed space, subgradients correspond to continuous linear functionals of dual norm \(1\) that align with the direction of \(x\). At \(x\neq 0\), \(\partial \|x\|\) typically has a structured set tied to the duality pairing.
  • Indicator functions: If \(f\) is the indicator of a convex set \(C\), meaning \(f(x)=0\) for \(x\in C\) and \(+\infty\) otherwise, then \(\partial f(x)\) is related to the normal cone of \(C\) at \(x\).
  • Piecewise-linear functions: At “kinks,” multiple affine pieces support the function. The subdifferential at such points contains all slopes of the active supporting pieces, often forming a finite interval (in one dimension) or a convex polytope (in higher dimensions).

4 Subdifferential calculus rules

4.1 Sum rules

A fundamental question is how \(\partial(f+g)\) relates to \(\partial f\) and \(\partial g\). Without assumptions, one typically obtains only inclusions. Under suitable regularity, the inclusion becomes an equality.

4.1.1 Conditions for exact sum of subdifferentials

A common exactness condition is that one function is well-behaved at a point where the other is finite, often expressed using interior/relative interior conditions on \(\mathrm{dom}\,f\) and \(\mathrm{dom}\,g\). Under such hypotheses, subgradients of the sum can be represented as sums of subgradients of the individual functions: \[ \partial(f+g)(\bar x)=\partial f(\bar x)+\partial g(\bar x) \] in the set-valued sense (Minkowski sum). When the regularity fails, only \[ \partial(f+g)(\bar x)\subseteq \partial f(\bar x)+\partial g(\bar x) \] may hold.

4.2 Chain rules

Chain rules adapt differentiation to nonsmooth settings by tracking how supporting hyperplanes transform through mappings.

4.2.1 Composition with affine maps

For an affine map \(A x+b\), subdifferentials obey a pullback relation. If \(h\) is convex and \(f(x)=h(Ax+b)\), then elements of \(\partial f(x)\) arise from pushing forward subgradients of \(h\) through the adjoint of \(A\). In finite dimensions, this reads in terms of transposes; in general spaces, it uses the appropriate adjoint operator between primal and dual spaces.

4.2.2 Composition with convex scalar functions

If \(f(x)=\phi(g(x))\) with \(\phi:\mathbb{R}\to\mathbb{R}\cup\{+\infty\}\) convex and \(g\) convex, the subdifferential involves both \(\partial g(x)\) and a scalar subgradient of \(\phi\) at \(g(x)\). Exactness can depend on qualification conditions ensuring that the chain rule is not weakened by nonsmooth interactions.

4.3 Product/maximum/minimum constructions

While multiplication is delicate in nonsmooth convex analysis, the maximum and related constructions are natural because convexity is preserved.

4.3.1 Subdifferential of a maximum

For convex functions \(f_i\) and \(f=\max_i f_i\), subgradients at a point typically come from the “active” indices: \[ \partial f(\bar x)=\mathrm{conv}\Bigl(\bigcup_{i\in I(\bar x)} \partial f_i(\bar x)\Bigr) \] under standard finite-maximum assumptions, where \(I(\bar x)=\{i:\ f_i(\bar x)=f(\bar x)\}\). Intuitively, at \(\bar x\) the maximum is realized by several supporting graphs, and the subdifferential averages the corresponding slopes.

4.3.2 Subdifferential of an indicator function

For the indicator \(\delta_C\) of a convex set \(C\), subgradients at \(x\in C\) form the normal cone \(N_C(x)\). When \(x\notin C\), the indicator is \(+\infty\), and the subdifferential is empty. This construction makes normal cones a central tool in constraint handling.

4.4 Linear change of variables

A linear transformation of variables modifies subgradients through dual mappings. If \(f(x)=h(Bx)\) with linear \(B\), then the subdifferential at \(x\) relates to \(\partial h(Bx)\) composed with \(B\) via the dual operator. These relationships are essential for reformulating constrained problems and for deriving dual formulations.

5 Subdifferential and optimality

5.1 First-order necessary and sufficient conditions

For convex optimization of the form \(\min_x f(x)\) with \(f\) proper convex, the subdifferential provides a complete first-order characterization. A point \(\bar x\in\mathrm{dom}\,f\) minimizes \(f\) if and only if \[ 0\in \partial f(\bar x). \] This criterion is both necessary and sufficient due to convexity: the supporting hyperplane inequality becomes a global statement that no direction can decrease the function.

5.2 KKT-type inclusions for convex problems

For problems that combine objectives and convex constraints, one often expresses optimality through inclusion relationships involving subdifferentials. For instance, in composite settings \(f(x)=\varphi(x)+\delta_C(x)\), the condition \(0\in \partial\varphi(\bar x)+N_C(\bar x)\) resembles Karush–Kuhn–Tucker (KKT) systems, though the exact structure depends on whether multipliers are represented explicitly or through normal cones and dual variables.

5.3 Variational inequalities and generalized gradients

Subdifferential conditions are closely tied to variational inequalities: the statement “\(\bar x\) is optimal” can be rewritten as an inequality against feasible directions, with subgradients playing the role of generalized gradients. This viewpoint is particularly useful in equilibrium and nonsmooth mechanics, where constraints define permissible moves and the optimality condition becomes a balance law.

5.4 Relationship to minimizers and critical points

In convex analysis, the terms “critical point” and “minimizer” coincide under mild assumptions when using the subdifferential: a critical point is typically defined by \(0\in \partial f(\bar x)\), which for convex \(f\) is exactly the minimizer condition.

5.4.1 Zero-subgradient criterion

The zero-subgradient criterion is a practical test: if one can compute or bound \(\partial f(\bar x)\), checking whether the origin belongs to the set decides optimality. In many algorithmic schemes, this becomes the stopping rule, often implemented indirectly via residuals related to monotone operators.

6 Connections to conjugate functions

6.1 Fenchel conjugate and duality basics

Given a proper convex function \(f:X\to\mathbb{R}\cup\{+\infty\}\), its Fenchel conjugate \(f^*:X^*\to\mathbb{R}\cup\{+\infty\}\) is defined by \[ f^*(\ell)=\sup_{x\in X}\bigl(\langle \ell,x\rangle - f(x)\bigr). \] Conjugacy links minimization in the primal to maximization in the dual. Under appropriate conditions, the optimal values agree (strong duality), and optimal solutions relate via subdifferentials.

6.2 Subdifferential of the conjugate

Subdifferentials of conjugates mirror each other. A core relation is that \(\ell\in \partial f(x)\) if and only if \(x\in \partial f^*(\ell)\), provided the relevant finiteness and regularity conditions hold.

6.2.1 Conjugacy optimality equivalences

The equality case of Fenchel inequalities yields practical equivalences: optimal primal-dual pairs correspond to mutual subgradient relations. This is the mechanism behind many duality theorems in convex optimization and variational analysis.

6.3 Moreau–Rockafellar-type relationships

Moreau–Rockafellar identities connect subdifferentials, conjugates, and normal cones, often expressing inversion of subgradients in terms of dual variables. Such results explain why conjugate functions are not merely transforms but encode geometric and variational information.

6.3.1 Inversion of subdifferentials

A typical inversion principle states that if \(\ell\in \partial f(x)\), then \(x\) lies in \(\partial f^*(\ell)\). When conditions are strengthened (e.g., under essential smoothness or strict convexity), these sets can become singletons, yielding an actual inverse mapping between gradients or subgradients.

6.4 Young’s inequality and equality cases

Fenchel–Young inequality asserts \[ f(x)+f^*(\ell)\ge \langle \ell,x\rangle. \] Equality holds precisely when \(\ell\in \partial f(x)\) (and equivalently \(x\in \partial f^*(\ell)\)). Thus, the subdifferential is the equality-enforcing locus for the primal–dual pairing inequality.

7 Examples and computations

7.1 Absolute value and piecewise linear functions

In one dimension, let \(f(x)=x\). At \(x\neq 0\), the function is differentiable and \(\partial f(x)=\{\mathrm{sign}(x)\}\). At \(x=0\), supporting slopes range from \(-1\) to \(+1\), giving \(\partial f(0)=[-1,1]\). This example illustrates the “interval of supporting slopes” phenomenon at a nonsmooth point.

7.2 Norms and dual norms

Consider \(f(x)=\|x\|\) with norm \(\|\cdot\|\) and dual norm \(\|\cdot\|_*\). For a given \(x\neq 0\), a functional \(\ell\) belongs to \(\partial \|x\|\) exactly when \(\|\ell\|_*=1\) and \(\langle \ell,x\rangle=\|x\|\). This characterization shows how subdifferentials encode the dual attainment of the norm.

7.3 Indicator functions and normal cones

For a closed convex set \(C\subset X\), the indicator \(\delta_C\) has subdifferential \[ \partial \delta_C(x)=N_C(x) \] for \(x\in C\). For many polyhedral sets in finite dimensions, normal cones are generated by outward normals of the active constraints, making \(\partial \delta_C(x)\) computable via linear inequalities describing \(C\).

7.4 Quadratic functions and affine shifts

Let \(f(x)=\frac{1}{2}\|x-a\|^2\) in a Euclidean space. The function is differentiable, so \(\partial f(x)=\{\nabla f(x)\}\), with

\[ \nabla f(x)=x-a. \] More generally, adding an affine term \( \langle b,x\rangle \) shifts the gradient by \(b\), and thus shifts the subdifferential by the same vector, reflecting the translation rule for affine perturbations.

7.5 Subdifferentials on polyhedral sets

For polyhedral convex functions (those built from maxima of affine functions or sums of linear and indicator terms), the subdifferential at a point is typically a polytope (convex hull of finitely many slopes). Computation reduces to identifying active pieces (or active constraints) and assembling the convex combination of their corresponding gradients or normal vectors.

8 Further theory and advanced topics

Normal cones describe how a convex set pushes back against perturbations. They connect subdifferentials of indicator functions to geometric properties such as curvature-like measures and to notions of generalized surface normals. These links are especially prominent in variational geometry, where sets of interest may have non-smooth boundaries.

8.2 Maximal monotone operator viewpoint

From the standpoint of monotone operator theory, \(\partial f\) is not only monotone but often maximal monotone under standard assumptions on \(f\). This viewpoint treats subdifferentials as operators whose resolvents and iterates can be studied for convergence and stability.

8.2.1 Graph and monotonicity interpretations

The monotonicity inequality corresponds exactly to a geometric non-crossing property of the graph of \(\partial f\). Maximality expresses that the graph cannot be enlarged while preserving monotonicity, providing a firm analytical foundation for algorithmic constructions like proximal mappings.

8.3 Legendre-type functions and differentiability

For strictly convex and essentially smooth convex functions (often grouped under Legendre-type conditions), subdifferentials behave like single-valued maps on the interior of the domain, and conjugacy yields strong differentiability correspondences. In such regimes, subdifferential calculus closely resembles classical differential calculus while preserving convexity and stability.

8.4 Regular subdifferentials versus general ones

Nonsmooth convex analysis has multiple subdifferential notions when moving beyond the basic convex subdifferential. “Regular” versions often correspond to supporting hyperplanes under stronger qualification constraints, while “general” versions permit weaker limits. Even in convex settings, different constructions can be used to handle degenerate or boundary behavior more robustly.

8.5 Limiting subdifferentials in nonsmooth extensions

For nonconvex or extended-value functions, one often defines limiting subdifferentials as outer limits of subgradients of nearby points or approximating functions. While these generalizations go beyond the basic convex theory, they preserve a similar role: subdifferentials become generalized gradients that still support first-order optimality conditions in variational models.