1 Basic definitions

Incidence relation is a mathematical expression used to connect incidence measures—typically derived from counts of new events—with foundational quantities such as the population size, the length of observation, and relevant risk-set definitions. In epidemiology and related fields, incidence relations formalize how often events occur, emphasizing that “new cases” must be defined relative to a reference state (for example, not already having the condition) and within a specified time span.

1.1 Incidence versus prevalence

Incidence describes the occurrence of new events in a population over a time interval, while prevalence describes the proportion of the population that has an event or condition at a point in time (or over a period). Incidence is therefore fundamentally about change and timing; prevalence is a stock-like measure influenced by both incidence and the duration of the condition. Because prevalence depends on how long cases persist, two settings can share the same incidence while exhibiting different prevalence if the duration of disease differs.

1.2 Population at risk

The population at risk is the subset of individuals who are eligible to experience the event during the interval under the study definition. This set typically excludes people who already had the event at baseline (for first-incidence questions) and may also exclude those under observation conditions where the event cannot be detected. The choice of risk set is central: the “denominator” in incidence relations should correspond to those truly capable of becoming incidents during the time window.

1.3 Time interval and counting conventions

Incidence relations require an explicit time frame, such as a fixed calendar period or an interval of follow-up measured per individual. Counting conventions determine whether events are attributed to the start of an interval, end of an interval, or treated as occurring at an approximate time within the interval. When events happen between observation times, discrete-time approximations and continuous-time models may yield different numerical results, especially when risks are large.

1.4 New cases (events) and censoring basics

A “new case” (event) is an occurrence that enters the incident count according to the study’s operational definition (e.g., first diagnosis, first hospitalization, first occurrence of a specified outcome). Censoring arises when an individual’s follow-up ends for reasons other than the event (loss to follow-up, administrative cutoff, or study end). In incidence rate relations, censoring is often handled by restricting the denominator to the time an individual remains under observation and at risk. In incidence proportion relations, censoring requires careful interpretation because not all individuals may contribute equal time.

2 Incidence proportion relations

Incidence proportion (also called cumulative incidence in many contexts) relates the number of new events over an interval to the number of individuals at risk at the beginning (or within the eligible baseline definition). It is a probability-like quantity, typically expressed as a proportion or percentage.

2.1 Formula for incidence proportion

A common form is \[ \text{Incidence proportion} = \frac{\text{Number of new events during the interval}}{\text{Number of individuals at risk at the start of the interval}}. \] This relation treats the denominator as the count of people who could have experienced the event during the same time interval. It implicitly assumes that each individual’s exposure opportunity aligns with the interval boundaries (or that censoring is negligible or handled by restricting to complete follow-up).

2.2 Discrete-time interpretation

In discrete-time settings (e.g., yearly follow-up), incidence proportion corresponds to the probability of experiencing the event within a time step, conditional on being event-free at the beginning of that step. When the event can occur at most once within the interval, the incidence proportion summarizes the chance of “becoming a case” during that interval. If multiple event occurrences are possible, incidence proportion may either focus on first events or require additional counting rules.

2.3 Complementary probability perspective

If the event is assumed to occur at most once in the interval, incidence proportion can be reframed using a complement: \[ \text{Incidence proportion} = 1 - \Pr(\text{no event in the interval}). \] This view highlights why incidence proportion is linked to survival-like quantities: the probability of remaining event-free over time is directly tied to how incidence accumulates.

2.4 Relation to risk and cumulative incidence

Incidence proportion is often used as an estimate of risk over a fixed window because it maps to a probability under standard assumptions. When the interval is longer, incidence proportion can be described as cumulative incidence: it aggregates the chance of experiencing the event up to that point. For higher risks, cumulative incidence grows more rapidly and becomes less symmetric with incidence rate, which spreads event occurrence over time using time at risk.

3 Incidence rate relations

Incidence rate expresses how rapidly events occur in a population, using person-time as the denominator. It is sensitive to differences in follow-up duration and is particularly useful when individuals contribute unequal time under observation.

3.1 Formula for incidence rate

A standard form is \[ \text{Incidence rate} = \frac{\text{Number of new events}}{\text{Total person-time at risk}}. \] “Person-time” aggregates the time each at-risk individual remains observable before experiencing the event or being censored.

3.2 Person-time and rate units

Person-time is typically measured in person-years, person-months, or person-days. If the numerator counts events, then the incidence rate has units of events per unit of person-time (for example, cases per person-year). Choosing consistent units is crucial for interpretability and for converting rates between studies that use different time scales.

3.3 Continuous-time interpretation

In continuous-time formulations, the incidence rate connects to the instantaneous occurrence of events. When follow-up is plentiful and event times are treated as continuous, rate-based measures can reflect how event intensity changes over time. This is especially relevant when risks vary across follow-up or when censoring is independent given covariates.

3.4 Relation to hazard under common assumptions

Under common modeling assumptions, the incidence rate can approximate a hazard (the instantaneous event rate among those still at risk). The hazard notion typically emerges in survival analysis, where event risk is expressed conditionally on survival (event-free status) up to a time point. While incidence rate and hazard are related, they are not identical in every data structure; the exact relationship depends on how events are distributed over time and whether the rate is averaged over a period or evaluated instantaneously.

4 Converting between incidence measures

Many analyses report incidence proportion while others use incidence rate. Conversion is possible under modeling assumptions that relate probability over time to rate (often through exponential or survival-style relationships).

4.1 Approximate conversions for small risks

When event probabilities over a short interval are small, the relationship between risk-like and rate-like measures can be approximated. For small risks, the incidence proportion over a window of length \( \Delta \) often behaves like \[ \text{Incidence proportion} \approx (\text{incidence rate}) \times \Delta. \] This approximation arises because higher-order terms (reflecting the growing chance of event occurrence as time accumulates) are small when the risk is low.

Conversion is frequently facilitated by survival-style functions. If the event-time distribution is modeled such that the survival probability declines exponentially with time at a constant hazard (a simplifying assumption), then the cumulative incidence over time \(t\) can be expressed as a complement of survival. Under those assumptions, rate-based quantities can be transformed into probability-like quantities and vice versa by using relationships of the form “probability of no event” versus “probability of event.”

4.3 Scaling across different time windows

When the hazard or rate is constant (or approximately constant), measures can be rescaled to other time windows. For instance, if a rate estimate is known and the rate is assumed stable, the probability over a longer window can be computed by applying the survival-style transformation repeatedly or directly via a closed-form expression. In real datasets, constancy may not hold, so scaling should be treated as an approximation unless supported by evidence.

4.4 Adjustments for varying follow-up

If individuals contribute different follow-up lengths, incidence proportion computed over a fixed interval may be distorted when censoring is substantial. In contrast, incidence rate naturally accommodates variable follow-up through person-time. Converting between measures typically requires either (a) information about the distribution of follow-up and censoring or (b) stronger assumptions that allow aggregation. Without these, conversions can misrepresent the underlying risk.

5 Role in statistical modeling

Incidence relations are not only descriptive; they also underpin many statistical procedures. Modeling uses the incidence-related quantities to connect data to parameters, often through likelihood or estimating-equation frameworks.

5.1 Incidence relations as likelihood components

In many statistical models, event counts and times enter through terms that reflect the incidence structure. For example, when counts arise over person-time, count models (or time-to-event models) use incidence-related denominators to produce a probability for the observed pattern of events and follow-up. The incidence relation thus provides the bridge from raw data to model-based inference.

5.2 Regression connections (high-level)

Regression modeling often expresses how incidence changes with predictors. At a high level, the regression framework selects a functional form that links covariates to incidence probability (incidence proportion) or to event intensity (incidence rate/hazard). The chosen form determines how effects combine over time and how interpretation should be mapped back to incidence quantities.

5.3 Interactions with time-dependent covariates

When covariates vary during follow-up (for example, changing exposure status), incidence relations must account for the evolving risk set and time-varying influence. This can be implemented by splitting follow-up into intervals where covariate values are constant or by using models designed for time-dependent predictors. The outcome is a more faithful representation of how incidence responds to changes over time.

5.4 Estimators and standard errors (conceptual)

Incidence relations inform which quantities are estimated (rates, probabilities, ratios) and how uncertainty is quantified. Standard errors depend on the data’s event structure—particularly the number of events, variability in follow-up, and censoring patterns. Conceptually, tighter confidence intervals arise when there are more events and more stable denominators (person-time or risk-set sizes), while sparse events lead to larger uncertainty.

6 Data requirements and practical considerations

Correct use of incidence relations depends on careful measurement of events, correct denominators, and appropriate handling of censoring and data quality issues.

6.1 How to define “incident” events

An incident event must be defined operationally and consistently. For first-incidence questions, baseline disease status should be assessed so that individuals already affected are excluded from the at-risk set. For outcomes that can recur, analysts must decide whether the incidence measure targets any recurrence, first recurrence only, or another event definition. Misalignment between the scientific question and the “incident” definition can produce misleading incidence estimates.

6.2 Dealing with left truncation and right censoring

Right censoring occurs when individuals leave follow-up without experiencing the event; it is handled via person-time contribution in rate-based approaches and via appropriate likelihood or weighting strategies in proportion-based approaches. Left truncation (delayed entry) occurs when individuals enter the study after beginning of the observation window. The at-risk set must then reflect eligibility from the time of entry forward, ensuring that denominators do not improperly include time before observation.

6.3 Handling repeat events (overview)

For repeat events, basic incidence relations must be extended or modified. Analysts may use methods that count only the first event, aligning with a probability interpretation, or they may model event processes that allow multiple events per person. The choice affects both the denominator and interpretation: counting all repeats changes what “incidence” represents, shifting from risk of first occurrence to event frequency.

6.4 Missingness and misclassification effects

Missing data can occur in event ascertainment, follow-up times, or covariates. If missingness is related to event likelihood and not properly addressed, incidence estimates may be biased. Misclassification of events (false positives/negatives) affects the numerator, while errors in determining who is at risk and for how long affect the denominator. Sensitivity analyses and validation studies are often used to evaluate the robustness of incidence relations to such imperfections.

7 Interpretation and limitations

Incidence relations are powerful, but their meaning depends on assumptions about the data-generating process and the study design.

7.1 Assumptions behind incidence relations

Key assumptions include correct identification of the at-risk population, consistent event definitions, and a coherent treatment of censoring. Rate-based interpretations often assume that censoring is non-informative conditional on what is modeled, meaning that those censored have the same future event intensity as those remaining at risk with similar covariate histories. Proportion-based measures additionally assume that the denominator corresponds to comparable follow-up opportunity or that censoring is negligible.

7.2 Common pitfalls in applying formulas

Pitfalls include using an incorrect denominator (such as including individuals who are not truly at risk), mixing time scales (e.g., numerator over one interval and denominator over another), and ignoring the impact of unequal follow-up when using incidence proportions. Another error is treating a proportion as a rate without accounting for time length, or rescaling measures across intervals without adequate justification of stability assumptions.

7.3 Comparing groups with different follow-up

When comparing groups that differ in follow-up duration or censoring patterns, crude incidence proportions can be misleading. Incidence rates often provide a more comparable summary because they incorporate person-time. When comparing rates, analysts may still need to consider whether hazards change over time differently across groups, which can affect interpretability of a single averaged rate ratio.

7.4 Communicating incidence results clearly

Clear communication requires specifying: (1) event definition, (2) time window, (3) risk-set definition, (4) whether the reported measure is a proportion or a rate, and (5) how censoring and follow-up were handled. Reporting units and denominators (risk set size or person-time) helps readers interpret magnitude and compare findings across studies.

8 Worked examples (statistics-focused)

These examples illustrate how incidence relations are computed and interpreted, emphasizing calculation steps and interpretation of ratios and conversions.

8.1 Example: incidence proportion calculation

Suppose a study follows \(N=1{,}000\) individuals who are event-free at baseline for one year. During that year, \(40\) individuals develop the event for the first time. The incidence proportion over the year is \[ \frac{40}{1{,}000} = 0.04, \] equivalently \(4\%\). If censoring is minimal and follow-up is complete for most participants, this can be interpreted as the risk of developing the event within one year.

8.2 Example: incidence rate from person-time

Consider another study with \(500\) individuals. Across follow-up, participants contribute a total of \(850\) person-years before experiencing the event or being censored. If \(34\) events occur, the incidence rate is \[ \frac{34}{850} = 0.04 \] events per person-year. Expressed per 1,000 person-years, this would be \(40\) events per 1,000 person-years. This rate accounts for differing follow-up times across individuals.

8.3 Example: interpreting ratios and differences

If group A has incidence proportion \(0.03\) and group B has \(0.06\), the risk ratio is \(0.06/0.03 = 2\), meaning the event probability in group B is twice that in group A over the same interval. The risk difference is \(0.06-0.03=0.03\), representing an absolute increase of 3 percentage points. For rate measures, an analogous comparison yields a rate ratio or a rate difference, using the respective rate denominators.

8.4 Example: converting measures across intervals

Assume a setting where the incidence rate is approximately constant at \(r=0.02\) events per person-year. Under an exponential (survival-style) approximation, the cumulative incidence over \(t=2\) years is \[ 1 - e^{-rt} = 1 - e^{-0.04} \approx 1 - 0.9608 = 0.0392, \] or about \(3.9\%\). For comparison, the “small risk” approximation would yield \(r \times t = 0.02 \times 2 = 0.04\), close to the exact value when the risks are modest.

Incidence relations connect to a family of measures used for summarizing event frequency and time-to-event processes.

9.1 Risk difference and risk ratio (connection)

Risk difference compares absolute event probabilities across groups, while risk ratio compares relative probabilities. Both are based on incidence proportion (cumulative risk) over a shared time window and are sensitive to that window’s length.

9.2 Rate ratio and rate difference (connection)

Rate ratio compares incidence rates computed from person-time across groups, capturing relative intensity of event occurrence. Rate difference summarizes absolute differences in intensity. These measures are often preferred when follow-up durations vary because person-time aligns denominators with exposure opportunities.

9.3 Survival analysis vocabulary

Terms such as survival probability, hazard, censoring, and risk set are central in survival analysis. Incidence relations for rates and cumulative incidence can be viewed through this vocabulary, particularly when incidence is framed as event-time behavior.

9.4 Further reading and standard notation guides

Standard references often provide consistent notation for: person-time denominators, hazard and survival functions, and transformations between probability-like and rate-like quantities. Familiarity with these conventions helps prevent unit errors and misinterpretation when applying incidence relations across studies and models.