1 Definition and basic concepts

The hazard function is a central concept in time-to-event analysis. It describes how quickly an event is occurring at a particular moment among individuals or items that have not yet experienced the event. In statistical and engineering settings, the event may be death, relapse, failure, or another outcome of interest.

1.1 Time-to-event framework

A time-to-event framework records the length of time until an event happens. The outcome is not simply whether the event occurs, but when it occurs. This perspective is useful when some observations are censored, meaning the event has not been observed by the end of follow-up or data collection.

1.2 Formal definition of the hazard function

The hazard function is defined for a nonnegative random variable representing event time. At a given time, it measures the instantaneous event rate conditional on survival up to that time. For continuous time, it is expressed as a limit involving a very small time interval.

1.2.1 Conditional event rate

The conditional event rate emphasizes that the hazard depends on having remained event-free until the time of interest. It is not a probability over a long interval, but a rate evaluated among the surviving population at that instant.

1.2.2 Instantaneous interpretation

The term instantaneous indicates that the hazard concerns an infinitesimally small interval around a time point. In practice, it is interpreted as the relative tendency for the event to occur immediately after that time, given that it has not occurred earlier.

1.3 Relation to survival and density functions

The hazard function is mathematically linked to other functions used to describe event times. These relationships make it possible to move between rate-based, probability-based, and cumulative descriptions of the same process.

1.3.1 Survival function

The survival function gives the probability that the event time exceeds a specified value. It summarizes the chance of remaining event-free up to that time and is directly connected to the hazard through an exponential relationship.

1.3.2 Probability density function

For continuous event times, the probability density function describes how event times are distributed across time. The hazard can be expressed using the density divided by the survival function, showing how the event intensity depends on both occurrence and persistence.

1.4 Cumulative hazard function

The cumulative hazard function is the accumulated hazard over time. It provides a running total of event intensity and is useful because it links the hazard and survival functions in a compact form. Larger cumulative hazard values generally correspond to lower survival probabilities.

2 Mathematical properties

The hazard function has several basic mathematical features that shape its use in modeling. Its form can vary greatly with time, which allows it to represent many kinds of event processes.

2.1 Nonnegativity

Hazard values are nonnegative because they represent rates of event occurrence. A negative hazard would not have a meaningful interpretation in standard survival or reliability settings.

2.2 Dependence on time

The hazard may change as time passes. This time dependence allows models to capture aging, wear, adaptation, or other dynamic patterns in risk.

2.2.1 Increasing hazard

An increasing hazard means the event becomes more likely as time advances, conditional on survival. This pattern is often associated with aging, deterioration, or progressive wear.

2.2.2 Decreasing hazard

A decreasing hazard indicates that the event rate declines over time among those still at risk. This may occur when early failures are more common or when susceptible cases are removed early from the population.

2.2.3 Constant hazard

A constant hazard remains unchanged over time. This is a simple and widely used assumption, particularly in the exponential model, where the event process has no duration dependence.

2.3 Hazard and probability models

Hazard functions appear in both continuous and discrete probability models. The meaning is similar across these settings, although the formulas differ because time is treated differently.

2.3.1 Continuous-time models

In continuous-time models, the hazard is defined through limiting behavior over very small intervals. This framework is common in medicine, engineering, and many statistical applications.

2.3.2 Discrete-time models

In discrete-time models, events are measured at specific intervals, such as days, months, or survey waves. The hazard is then interpreted as the conditional probability of event occurrence in a given interval.

2.4 Transformation relationships

The hazard, survival, and cumulative hazard functions can be transformed into one another. These transformations are fundamental for both theoretical derivations and practical estimation.

2.4.1 From hazard to survival

Given a hazard function, the survival function can be obtained by integrating the hazard over time and applying an exponential transformation. This shows how accumulated risk reduces the probability of remaining event-free.

2.4.2 From survival to hazard

If the survival function is known and sufficiently smooth, the hazard can be recovered from it. This is done by differentiating the survival function in a way that isolates the instantaneous rate of event occurrence.

3 Types of hazard functions

Different forms of hazard functions are used to reflect distinct scientific questions and study designs. The terminology often depends on whether one is studying a general baseline process, competing events, or individual-specific conditions.

3.1 Baseline hazard

The baseline hazard is the underlying hazard in a reference group or under baseline covariate values. It serves as the starting point for models that adjust the hazard according to predictors or group membership.

3.2 Cause-specific hazard

A cause-specific hazard describes the instantaneous rate of a particular event type in the presence of other possible event types. It is commonly used in competing risks analysis, where several mutually exclusive outcomes may occur.

3.3 Subdistribution hazard

The subdistribution hazard is another competing-risks measure that focuses on cumulative incidence for a particular event type. It is used when direct modeling of the probability of a specific outcome over time is the main goal.

3.4 Conditional hazard

A conditional hazard is the hazard given additional information beyond survival up to time alone. This may include covariates, previous events, or other subject-specific characteristics that alter the risk pattern.

4 Applications

Hazard functions are widely applied because they provide a flexible way to study timing. They help analysts compare risks, estimate event probabilities, and build models that reflect changing conditions over time.

4.1 Survival analysis

In survival analysis, the hazard function is a core quantity for studying durations until an event occurs. It helps describe prognosis, assess treatment effects, and handle censored observations.

4.1.1 Medical prognosis

In medical settings, hazards are used to characterize the risk of death, recurrence, or complication over time. They support the analysis of patient trajectories and the comparison of outcomes across clinical groups.

4.1.2 Clinical trials

Clinical trials often use hazard-based methods to compare treatment arms. These methods can summarize differences in event timing rather than only overall event counts.

4.2 Reliability engineering

In reliability engineering, the hazard function is often called the failure rate. It describes the chance that a component or system fails at a given time after having functioned up to that point.

4.2.1 Component failure

For individual parts, the hazard helps identify whether failures are more likely early in life, during a stable operating period, or as a result of aging. This is useful in testing, maintenance, and quality control.

4.2.2 System lifetime analysis

For complete systems, hazard analysis can reveal how the arrangement of components affects overall lifetime. It is used to design systems with longer operating periods and fewer unexpected breakdowns.

4.3 Demography and actuarial science

In demography, hazard functions are used to study mortality and other life course events. In actuarial science, they support calculations related to life expectancy, insurance risk, and benefit planning.

4.4 Economics and social science

Economists and social scientists use hazards to examine unemployment duration, job transition, marriage timing, and other social processes. The framework is valuable for studying how risk changes with age, experience, or other covariates.

5 Estimation and inference

Because the hazard is rarely observed directly, it must usually be estimated from data. Several approaches are available, ranging from simple nonparametric methods to flexible regression models.

5.1 Nonparametric estimation

Nonparametric methods estimate event-time behavior without imposing a specific distributional form. They are often used for descriptive analysis and as a foundation for more advanced modeling.

5.1.1 Kaplan–Meier-based methods

Kaplan–Meier-based methods estimate survival first and then infer hazard-related quantities from it. They are especially useful when data include censoring and the analyst wants a direct summary of observed event experience.

5.1.2 Nelson–Aalen estimator

The Nelson–Aalen estimator provides a nonparametric estimate of cumulative hazard. It is based on observed event counts and the number at risk over time, making it a standard tool in survival analysis.

5.2 Parametric models

Parametric models assume a particular mathematical form for the hazard. These models are useful when the assumed shape is plausible and when a compact description of the event process is desired.

5.2.1 Exponential model

The exponential model assumes a constant hazard over time. It is mathematically simple and often serves as a baseline case for more flexible models.

5.2.2 Weibull model

The Weibull model allows the hazard to increase or decrease with time depending on its parameters. This flexibility makes it one of the most widely used parametric survival models.

5.2.3 Gompertz model

The Gompertz model is often used for adult mortality and other processes where risk rises approximately exponentially with time. It provides a convenient representation of steadily increasing hazard.

5.3 Semiparametric models

Semiparametric models combine a flexible baseline hazard with structured effects of explanatory variables. They are especially useful when the exact shape of the hazard is unknown but covariate effects are of interest.

5.3.1 Cox proportional hazards model

The Cox proportional hazards model is one of the most influential methods in survival analysis. It estimates relative effects of covariates on the hazard while leaving the baseline hazard unspecified.

5.3.2 Time-dependent covariates

Time-dependent covariates are variables whose values change during follow-up. Including them allows the hazard to respond to evolving conditions such as treatment changes, aging, or exposure history.

5.4 Model assessment

Model assessment checks whether the chosen hazard model fits the data adequately. It also evaluates whether the model assumptions are reasonable for the problem at hand.

5.4.1 Goodness of fit

Goodness of fit methods compare model predictions with observed event patterns. They help determine whether a proposed hazard structure captures the main features of the data.

5.4.2 Residual diagnostics

Residual diagnostics examine departures between observed and fitted values. In survival models, these tools can reveal misspecification, unusual observations, or violations of model assumptions.

6 Interpretation and usage

Interpreting hazard functions requires attention to both the underlying rate and the conditioning on survival. The hazard is useful for comparison, but it should not be confused with a direct probability of event occurrence over a long period.

6.1 Comparing hazards across groups

Comparing hazards across groups can reveal whether one group experiences events earlier or more frequently at specific times. Such comparisons are common in medical, engineering, and social research.

6.2 Hazard ratios

A hazard ratio compares the hazard in one group with the hazard in another. A value greater than one indicates higher instantaneous risk in the numerator group, while a value less than one indicates lower risk.

6.3 Assumptions and limitations

Hazard-based analyses often depend on assumptions about censoring, independence, proportionality, or the form of the event-time process. If these assumptions fail, interpretation may become unreliable or require alternative methods.

6.4 Common misconceptions

A frequent misconception is that the hazard is the same as the probability of event occurrence over a time interval. In fact, it is a rate conditional on survival and may not correspond directly to an observed probability unless converted through the survival function.

Several closely related terms are used alongside the hazard function. These concepts overlap in meaning but are not identical, and each emphasizes a different aspect of time-to-event behavior.

7.1 Survival function

The survival function gives the probability of remaining event-free beyond a given time. It is the most direct complement to the hazard in continuous-time analysis.

7.2 Failure rate

The failure rate is a common engineering term for the hazard function. It refers to the instantaneous tendency for a component or system to fail.

7.3 Mean residual life

Mean residual life is the expected remaining time until the event, given survival to the current time. It complements the hazard by focusing on remaining duration rather than instantaneous risk.

7.4 Risk function

Risk function is a broader expression that can refer to the likelihood or intensity of an event under a given model. In some contexts, it is used interchangeably with hazard, though usage may vary by discipline.