1 Concept and purpose
Effect size is a numerical summary of the magnitude of a result. It can describe the strength of a difference between groups, the degree of association between variables, or the size of a treatment or exposure effect. In statistical practice, it complements hypothesis tests by addressing not just whether an effect is detectable, but how large it is.
1.1 Definition of effect size
An effect size is any statistic that quantifies the size of an observed relationship or difference. Some effect sizes are expressed in the original measurement units, while others are standardized to make results comparable across studies or scales. The most suitable measure depends on the research question, data type, and analysis method.
1.2 Effect size versus statistical significance
Statistical significance indicates whether data are inconsistent with a null hypothesis under a chosen threshold, such as a p-value cutoff. Effect size, by contrast, estimates the magnitude of the observed pattern. A small effect can be statistically significant in a very large sample, while a large effect may fail to reach significance in a small sample.
1.3 Why effect size matters
Effect size is important because it helps researchers interpret results in substantive terms. It can show whether a finding is likely to have practical value, support comparison among studies using different scales, and provide a common basis for combining evidence.
1.3.1 Practical significance
Practical significance concerns whether an effect is large enough to matter in real-world settings. A statistically detectable change may still be too small to influence decisions, outcomes, or behavior in a meaningful way. Effect size helps distinguish trivial findings from those with substantive impact.
1.3.2 Comparison across studies
Studies often use different instruments, units, or measurement ranges. Standardized effect sizes reduce these differences, making it easier to compare results across experiments or observational analyses. This is especially useful when outcomes are measured with distinct scales.
1.3.3 Role in meta-analysis
Meta-analysis relies heavily on effect sizes because it combines evidence from multiple studies. Effect size metrics provide a common scale on which results can be synthesized, weighted, and compared. They also support assessment of variability among studies.
2 Types of effect size measures
Effect size measures vary according to the kind of data and statistical model involved. Some summarize mean differences, others describe correlation or association, and others are tied to analysis of variance or regression models. Selecting an appropriate metric depends on the structure of the data.
2.1 Standardized mean differences
Standardized mean differences express the size of a group difference relative to the variability in the data. They are commonly used when comparing two means measured on the same or similar scales.
2.1.1 Cohen’s d
Cohen’s d is one of the most widely used standardized mean difference measures. It is calculated as the difference between two group means divided by a pooled standard deviation. The result gives a scale-free estimate of how far apart the groups are.
2.1.2 Hedges’ g
Hedges’ g is similar to Cohen’s d but includes a small-sample correction. It is often preferred when sample sizes are limited, since it reduces bias in the estimated standardized difference.
2.1.3 Glass’s Δ
Glass’s Δ uses the standard deviation of the control group rather than a pooled standard deviation. It is useful when the treatment changes variability or when the control group provides a more stable reference point.
2.2 Correlation-based measures
Correlation-based effect sizes describe the degree and direction of association between variables. They are often used when the interest is in how strongly two measures move together.
2.2.1 Pearson’s r
Pearson’s r measures the strength of a linear relationship between two continuous variables. Values range from -1 to 1, with the sign indicating direction and the absolute value indicating strength.
2.2.2 Rank-based correlations
Rank-based correlations, such as Spearman’s rho and Kendall’s tau, assess monotonic association using the ordering of values rather than their raw distances. They are useful when data are ordinal, non-normally distributed, or influenced by outliers.
2.2.3 Coefficient of determination
The coefficient of determination, often written as r-squared in simple settings, indicates the proportion of variance in one variable explained by another. It is derived from correlation or regression models and offers an intuitive measure of explained variation.
2.3 Measures for categorical data
Categorical data often require effect size measures that compare proportions, risks, or odds. These metrics are common in biomedical and observational research.
2.3.1 Risk difference
Risk difference is the absolute difference in probability between two groups. It directly shows how much the event rate changes, making it easy to interpret in applied contexts.
2.3.2 Relative risk
Relative risk compares the probability of an outcome in one group with that in another. A value above 1 indicates increased risk, while a value below 1 indicates reduced risk relative to the reference group.
2.3.3 Odds ratio
The odds ratio compares the odds of an event between two groups. It is widely used in case-control studies and logistic regression, although it is often less intuitive than risk-based measures.
2.4 Measures for analysis of variance
In analysis of variance, effect sizes summarize how much variability in an outcome is associated with group membership or experimental factors.
2.4.1 Eta squared
Eta squared represents the proportion of total variance in the outcome explained by a factor. It is frequently reported in studies using analysis of variance.
2.4.2 Partial eta squared
Partial eta squared estimates the proportion of variance explained by a factor after accounting for other factors in the model. It is common in multifactor designs, though it is not always directly comparable across studies.
2.4.3 Omega squared
Omega squared is an alternative variance-explained measure that adjusts for bias more conservatively than eta squared. It is often considered a more realistic estimate of effect magnitude in some settings.
2.5 Measures for model fit and association
Model-based effect sizes summarize how well a statistical model accounts for variation or how strongly variables are associated within a fitted framework.
2.5.1 R-squared
R-squared indicates the proportion of variance in an outcome explained by a regression model. Higher values suggest greater explanatory power, though interpretation depends on the domain and the complexity of the data.
2.5.2 Phi coefficient
The phi coefficient measures association between two binary variables. It is especially useful for 2-by-2 contingency tables.
2.5.3 Cramér’s V
Cramér’s V generalizes association measures to larger contingency tables. It provides a normalized index of relationship strength for categorical variables with more than two levels.
3 Interpretation
Interpreting effect size requires attention to context, measurement, and research goals. A value that appears small in one field may be substantial in another, and the same numerical effect can have different practical implications depending on the situation.
3.1 Conventional benchmarks
Researchers sometimes use conventional cutoffs to describe effect sizes as small, medium, or large. These labels can be helpful as rough guides, but they should not be treated as universal standards.
3.1.1 Small, medium, and large labels
Common benchmarks have been proposed for several effect size metrics, especially standardized mean differences and correlations. These categories provide a shorthand for discussion, but they do not replace substantive interpretation.
3.1.2 Field-specific standards
Different disciplines develop their own expectations based on typical data patterns and substantive importance. In some areas, modest numerical effects can be highly consequential, while in others larger estimates are needed to be meaningful.
3.2 Context dependence
The interpretation of an effect size depends on study design, measurement quality, and the underlying phenomenon. A single numerical value rarely has the same significance in every setting.
3.2.1 Sample size considerations
Sample size affects the precision of effect size estimates, not the underlying meaning of the metric itself. Large samples can produce very stable estimates, while small samples may yield more variable results.
3.2.2 Measurement scale considerations
Effect sizes may be influenced by how variables are measured. Less reliable or more restricted measures can attenuate observed effects, whereas well-calibrated instruments may yield stronger and clearer estimates.
3.2.3 Practical importance
Practical importance asks whether the effect changes decisions, outcomes, or interpretation in a meaningful way. This consideration is especially relevant in applied research, where numerical magnitude must be linked to real consequences.
4 Estimation and reporting
Effect sizes should be estimated carefully and reported in a way that allows others to interpret and reuse the results. Clear reporting improves transparency and supports comparison across studies.
4.1 Point estimates
A point estimate gives the single best numerical summary of the effect in a sample. It is typically reported alongside the method used to compute it and the comparison or model from which it was derived.
4.2 Confidence intervals
Confidence intervals show a range of plausible values for the effect size. They convey both magnitude and uncertainty, making them more informative than a point estimate alone.
4.3 Standard errors and precision
Standard errors describe the sampling variability of an effect size estimate. Smaller standard errors indicate greater precision, while larger ones suggest more uncertainty about the true effect.
4.4 Reporting in research articles
Research articles often report the effect size, its confidence interval, and the relevant statistical test. Good reporting also identifies the exact metric used, the groups or variables compared, and any transformations or corrections applied.
5 Effect sizes in study design
Effect sizes play an important role before data are collected. They inform how much data are needed, how likely a study is to detect an effect, and how sensitive the design is to different magnitudes.
5.1 Power analysis
Power analysis uses an assumed effect size to estimate the probability of detecting an effect if it exists. It helps align study design with the expected strength of the phenomenon under investigation.
5.2 Sample size planning
Sample size planning often begins with a target effect size. Researchers use this information to determine how many observations are needed for a study to be adequately informative.
5.3 Sensitivity analysis
Sensitivity analysis asks what effect sizes a study could detect given its sample size and design. This approach helps clarify the limits of a study’s evidential capacity.
6 Effect sizes in meta-analysis
Meta-analysis depends on effect sizes as a common currency for combining findings. It uses these estimates to summarize evidence across studies and examine patterns of consistency or variation.
6.1 Converting between effect size metrics
Different studies may report different statistics, so conversions are sometimes needed. Standard formulas can translate among several common metrics, such as correlations, standardized mean differences, and odds ratios.
6.2 Weighting and pooling estimates
Meta-analytic methods often weight effect sizes by precision, giving more influence to more reliable estimates. The pooled result then represents an overall summary of the available evidence.
6.3 Heterogeneity and variability
Heterogeneity refers to differences among study effect sizes beyond chance. It may arise from differences in participants, measurements, interventions, or settings, and it is a central concern in meta-analysis.
7 Limitations and common misconceptions
Effect sizes are valuable, but they are not self-explanatory. Misuse can lead to overinterpretation, false comparisons, or confusion about what a metric actually represents.
7.1 Misinterpretation of magnitude
A numerical value does not automatically indicate importance. Effect sizes should be interpreted in relation to baseline rates, measurement reliability, and the practical context of the finding.
7.2 Dependence on study design
Some effect sizes depend on the design, comparison group, or model specification. As a result, identical numerical values may not correspond to the same substantive meaning across different analyses.
7.3 Noncomparability across metrics
Not all effect size measures are directly comparable. A value from one metric cannot always be meaningfully contrasted with a value from another without careful conversion or contextual explanation.
8 Applications
Effect size is used across many quantitative disciplines because it provides a common way to describe the magnitude of empirical findings. Its exact form varies by field, but the general purpose remains the same.
8.1 Psychology and behavioral science
In psychology, effect sizes are used to summarize treatment effects, correlations among traits, and differences between experimental conditions. They are central to experimental reports and meta-analytic reviews.
8.2 Medicine and clinical research
Clinical research often uses effect sizes to compare outcomes between treatment and control groups, especially in trials and observational studies. Risk-based measures and standardized differences are both common.
8.3 Education and social science
Education research frequently reports effect sizes for interventions, achievement differences, and program evaluations. Social science studies also use them to compare outcomes across groups or policy settings.
8.4 Economics and other quantitative fields
Economics and related fields use effect sizes in regression, causal inference, and model comparison. These measures help quantify how strongly predictors are associated with outcomes and how much variation a model explains.