1 Fundamental concepts
1.1 Definition and purpose
A statistical hypothesis is a formal statement about a population, a probability distribution, or a model parameter. It is framed so that data can be used to assess whether the statement is plausible. In practice, hypotheses provide a structured way to translate a question into a form that can be examined with statistical methods.
The main purpose of a statistical hypothesis is to support inference under uncertainty. Rather than proving a claim outright, it allows analysts to compare observed evidence with what would be expected if the claim were true. This makes hypothesis formulation central to many quantitative studies.
1.2 Populations, samples, and parameters
Statistical hypotheses typically refer to a population, which is the full set of individuals, events, or measurements of interest. Because examining an entire population is often impractical, researchers usually work with a sample drawn from it. The sample is then used to learn about unknown population characteristics.
These unknown characteristics are called parameters. Examples include a population mean, a proportion, a variance, or a regression coefficient. A hypothesis may state a specific value for a parameter, a range of possible values, or a relationship among several parameters.
1.3 Role in statistical inference
Hypotheses are a foundation of statistical inference, the process of drawing conclusions from data beyond the observed sample. Inference links sample results to broader claims about a population or model. Hypothesis testing is one of the most widely used inferential procedures.
In this framework, data are evaluated against a proposed statement using probability-based reasoning. The resulting conclusion is not certainty, but a measured assessment of whether the evidence is consistent with the hypothesis. This approach is widely used because it imposes clear decision rules and makes uncertainty explicit.
1.4 Hypotheses in scientific research
In scientific research, hypotheses help define predictions, organize experiments, and connect theory to observation. A research question is often converted into a testable statement before data collection begins. This improves clarity and reduces ambiguity in interpretation.
Hypotheses may arise in laboratory studies, surveys, field observations, and observational research. They can address whether a treatment changes an outcome, whether two groups differ, or whether variables are associated. In many fields, careful hypothesis formulation is considered essential for reproducible and interpretable results.
2 Types of statistical hypotheses
2.1 Null hypothesis
The null hypothesis is the default statement tested in standard procedures. It commonly expresses no effect, no difference, or no association. For example, it may claim that a population mean equals a specified value or that two groups have the same average outcome.
The null hypothesis is not assumed to be true in a philosophical sense; rather, it serves as a reference point. Statistical evidence is then examined to determine whether the sample provides enough reason to doubt it. This convention supports consistent testing across many settings.
2.2 Alternative hypothesis
The alternative hypothesis is the competing statement that differs from the null. It represents the presence of an effect, difference, or association. Depending on the context, it may specify only that a departure exists or it may indicate a direction.
The alternative is the conclusion supported when the null is rejected. It does not necessarily prove a precise causal mechanism, but it indicates that the observed data are unlikely under the null model alone. The form of the alternative strongly influences the test procedure.
2.2.1 One-sided alternatives
A one-sided alternative specifies a direction, such as greater than or less than. For instance, a researcher may hypothesize that a new method increases performance relative to an established benchmark. This type of hypothesis is useful when only one direction is of substantive interest.
One-sided tests can be more powerful in the chosen direction than two-sided tests. However, they require a pre-established rationale, since the analysis focuses on one tail of the distribution. They are not appropriate when deviations in either direction matter equally.
2.2.2 Two-sided alternatives
A two-sided alternative states that a parameter differs from a specified value without indicating direction. It is commonly used when both increases and decreases are relevant. For example, a study may ask whether a mean differs from a target value, regardless of whether it is higher or lower.
Two-sided formulations are often the default in general statistical practice. They are especially appropriate when researchers do not have a strong basis for predicting direction. This makes them more flexible, though sometimes less sensitive to directional departures than one-sided tests.
2.3 Simple and composite hypotheses
A simple hypothesis specifies a single exact value for a parameter or a fully determined distribution. For example, a claim that a population mean equals 10 is simple if all other aspects of the model are fixed. Such hypotheses are mathematically precise and often easier to analyze.
A composite hypothesis allows multiple possible values. A statement that a mean is greater than 10, or that a parameter lies within a range, is composite. Composite forms are common in practice because many real questions concern sets of values rather than a single exact point.
2.4 Point and interval hypotheses
A point hypothesis assigns one exact value to a parameter. This is the most specific form and is often used for null hypotheses. It may state, for example, that a proportion equals 0.50 or that a variance equals a particular number.
An interval hypothesis allows values within a specified range. It is useful when the question concerns acceptable limits rather than a precise number. Interval formulations appear in equivalence testing, tolerance assessment, and other contexts where practical thresholds matter.
3 Formulating hypotheses
3.1 Research question to statistical statement
Formulating a hypothesis usually begins with a substantive research question. The question is then translated into a statistical claim that can be checked with data. This translation step requires defining the population, outcome, and comparison of interest.
A well-formed hypothesis is clear, measurable, and tied to an observable quantity. Vague statements are difficult to test, while precise statements allow the appropriate model and test statistic to be chosen. Good formulation often determines how meaningful the analysis will be.
3.2 Identifying variables and parameters
The next step is to identify the relevant variables and parameters. The variables may include group membership, treatment status, response values, or predictor measurements. The parameter is the numerical feature that summarizes the pattern in the population.
For example, a comparison of two groups may involve the difference between their population means. A survey question may involve a population proportion. More complex studies may involve regression coefficients, odds ratios, or correlation measures.
3.3 Choosing the direction of the test
Choosing whether to use a one-sided or two-sided test is an important part of hypothesis specification. The choice should be made before seeing the data, based on the research objective and the consequences of different outcomes. Post hoc selection can distort interpretation.
A directional test is most appropriate when only one direction has practical relevance and the opposite direction would not support the intended claim. If both directions matter, a two-sided hypothesis is usually preferred. The decision affects the allocation of error probabilities and the interpretation of results.
3.4 Examples of hypothesis statements
A common example is a test of whether the mean height of a population equals 170 cm. The null hypothesis states that the mean is 170 cm, while the alternative may state that it differs from 170 cm. This is a typical two-sided setup.
Another example concerns proportions: a manufacturer may want to test whether the fraction of defective items exceeds a threshold. In that case, the null might state that the defect rate is at most the threshold, while the alternative states that it is greater. Similar structures are used for variance, correlation, and regression hypotheses.
4 Hypothesis testing framework
4.1 Test statistics
A test statistic is a numerical summary computed from the sample. It measures how far the observed data are from what the null hypothesis predicts. Common test statistics include standardized differences, ratios, and chi-square measures.
The statistic is chosen so that its sampling distribution is known or can be approximated under the null hypothesis. This allows the observed value to be compared with its expected behavior. The farther the statistic falls into the tail of the null distribution, the stronger the evidence against the null.
4.2 Significance level
The significance level, usually denoted by alpha, is a preselected threshold for deciding whether evidence is strong enough to reject the null hypothesis. It represents the maximum tolerated probability of a Type I error under the testing rule. Common choices include 0.05 and 0.01.
The significance level is not the probability that the null is true. Instead, it is a property of the procedure, not of the specific conclusion. Choosing a smaller alpha makes rejection more difficult, while a larger alpha makes the test more permissive.
4.3 P-values
A p-value is the probability, under the null hypothesis, of observing a result at least as extreme as the one obtained. It summarizes how surprising the data would be if the null were correct. Smaller values indicate less compatibility with the null model.
A p-value is often misinterpreted, so its meaning must be stated carefully. It does not give the probability that the hypothesis is true or false. Rather, it measures extremeness relative to the null distribution and the chosen test statistic.
4.4 Critical regions
A critical region is the set of test statistic values that lead to rejection of the null hypothesis. Its boundaries are determined by the significance level and the distribution of the test statistic. Observations falling into this region are considered sufficiently unusual under the null.
Critical regions provide a direct decision rule. In a two-sided test, they are typically located in both tails of the distribution. In a one-sided test, the region lies in only one tail, corresponding to the chosen direction.
4.5 Decision rules
Decision rules specify what action follows from the observed test result. The simplest rule is to reject the null if the p-value is less than or equal to alpha, and otherwise fail to reject it. Equivalent rules can be written in terms of critical values.
These rules are designed to make testing reproducible and objective. They also help separate the statistical procedure from later substantive interpretation. A formal decision, however, should still be considered alongside study design, effect size, and assumptions.
5 Common test procedures
5.1 Tests for means
Tests for means evaluate hypotheses about population averages. They are among the most familiar procedures in statistics and appear in many scientific applications. The chosen method depends on whether one or more samples are involved and on what assumptions are reasonable.
5.1.1 One-sample tests
A one-sample mean test compares a sample average with a hypothesized population mean. It is often implemented with a t-test when the population variance is unknown and the sample is not extremely large. The procedure assesses whether the observed mean differs materially from the reference value.
This test is useful when a process is compared with a target or standard. Examples include quality control, baseline assessment, and performance benchmarking. The underlying logic is straightforward: determine whether the sample mean is too far from the claimed mean to be explained by random variation alone.
5.1.2 Two-sample tests
Two-sample mean tests compare the averages of two groups. They may be used to evaluate treatment effects, group differences, or changes over time. Variants include methods that assume equal variances and methods that do not.
A two-sample test asks whether the observed difference between group means is larger than would be expected from sampling fluctuation. This kind of test is common in experiments and comparative studies. Proper design and independent sampling are important for valid interpretation.
5.2 Tests for proportions
Tests for proportions assess claims about the share of cases with a particular attribute. They are used when outcomes are binary, such as success or failure, present or absent, or yes or no. A proportion test may compare a sample proportion with a reference value or compare proportions across groups.
These tests are often based on binomial or normal approximations, depending on sample size. They are widely used in surveys, medicine, manufacturing, and social research. The interpretation is similar to mean testing, but the parameter of interest is a probability rather than a numerical average.
5.3 Tests for association and independence
Tests for association and independence examine whether two variables are related. In contingency tables, a chi-square test of independence is commonly used to determine whether observed counts depart from what would be expected if the variables were unrelated. Correlation and regression tests serve similar purposes in other settings.
These procedures are helpful for categorical and continuous data alike. They do not automatically establish causation, but they can identify patterns worth further study. Their results depend on the chosen model and the quality of measurement.
5.4 Goodness-of-fit tests
Goodness-of-fit tests evaluate whether a sample follows a specified distribution. They compare observed data with expected frequencies or probabilities under the hypothesized model. Common examples include the chi-square goodness-of-fit test and tests based on empirical distribution functions.
Such tests are important when assessing whether a distributional assumption is reasonable. They may be used to check whether counts follow a theoretical pattern or whether a continuous variable is consistent with a proposed distribution. A poor fit can indicate model misspecification or unusual data structure.
5.5 Tests for variance and distributional assumptions
Tests for variance examine whether the variability in a population matches a hypothesized level. These are useful in settings where spread is as important as central tendency. They can also support comparisons of variability across groups.
Other procedures evaluate distributional assumptions such as normality or independence. These checks matter because many methods rely on approximations that work best when assumptions are reasonably satisfied. When assumptions fail, alternative models or robust methods may be more appropriate.
6 Interpretation and inference
6.1 Rejecting and failing to reject the null hypothesis
Rejecting the null hypothesis means that the data provide sufficient evidence against it at the chosen significance level. Failing to reject the null means that the evidence is not strong enough to justify rejection. It does not mean the null has been proven true.
This distinction is essential. A nonsignificant result may arise because the null is correct, because the effect is small, or because the study lacks precision. Interpretation should therefore focus on what the data do and do not support, rather than treating failure to reject as confirmation.
6.2 Statistical significance versus practical significance
Statistical significance refers to whether an observed result is unlikely under the null hypothesis. Practical significance concerns whether the size of the effect is large enough to matter in real-world terms. The two are related but not identical.
Large samples can produce statistically significant results even when effects are very small. Conversely, a meaningful effect may not reach statistical significance in a small or noisy study. Good analysis considers both the strength of evidence and the substantive importance of the estimate.
6.3 Confidence intervals and hypothesis testing
Confidence intervals and hypothesis tests are closely connected. A confidence interval gives a range of plausible values for a parameter, while a hypothesis test evaluates a specific claim about that parameter. In many standard settings, a null value is rejected if it lies outside the corresponding confidence interval.
Confidence intervals add information by showing the estimated magnitude and precision of the effect. They often help analysts interpret results more fully than a p-value alone. When used together, the two methods provide complementary perspectives on uncertainty.
6.4 Effect size and uncertainty
Effect size measures the magnitude of a difference, association, or departure from the null. It is useful because statistical significance alone does not reveal how large or meaningful the observed pattern is. Common effect measures include mean differences, risk ratios, correlations, and standardized differences.
Uncertainty describes how much the estimate might vary from sample to sample. It is usually summarized by standard errors, confidence intervals, or posterior distributions in other frameworks. Good inference balances effect size with uncertainty rather than focusing on a single threshold.
7 Assumptions and limitations
7.1 Model assumptions
Hypothesis tests usually rely on assumptions about the sampling process or underlying model. These may include random sampling, independence, equal variances, or approximate normality. When the assumptions are reasonable, the test result is more trustworthy.
If assumptions are badly violated, the nominal error rates may no longer hold. Analysts may then need robust methods, transformations, exact tests, or revised models. Assumption checking is therefore part of responsible statistical practice.
7.2 Sampling error and variability
Sampling error is the difference between a sample statistic and the corresponding population quantity caused by random selection. It is unavoidable in finite samples and is the reason hypothesis testing is needed. Variability also arises from measurement noise and natural heterogeneity.
Greater sample sizes generally reduce uncertainty, though they do not eliminate it. Recognizing variability prevents overinterpretation of small fluctuations. Hypothesis tests are designed to distinguish ordinary sample variation from unusually large deviations.
7.3 Type I and Type II errors
A Type I error occurs when the null hypothesis is rejected even though it is true. A Type II error occurs when the null is not rejected even though the alternative is true. These two error types reflect different kinds of testing mistakes.
The chosen significance level controls the long-run rate of Type I errors under the model. Type II errors are influenced by sample size, variability, effect magnitude, and the selected alpha level. Since reducing one error can increase the other, test design involves trade-offs.
7.3.1 Power of a test
Power is the probability that a test correctly rejects the null when a specified alternative is true. It depends on the true effect size, sample size, noise level, and significance threshold. Higher power means a lower chance of missing a real effect.
Researchers often use power analysis when planning studies. Adequate power improves the likelihood of detecting meaningful effects and reduces ambiguous findings. It is especially important when data collection is expensive or difficult.
7.3.2 False positives and false negatives
A false positive is a result that suggests an effect when none exists, corresponding to a Type I error. A false negative is a missed detection of a real effect, corresponding to a Type II error. These terms are often used in applied fields because they are intuitive.
The balance between them depends on the testing context. In some settings, avoiding false positives is crucial; in others, missing a genuine effect is more costly. The best choice of testing strategy depends on the consequences of each kind of mistake.
7.4 Multiple testing and selection bias
Multiple testing occurs when many hypotheses are examined at once. As the number of tests increases, the chance of at least one false positive also rises unless adjustments are made. Common remedies include p-value correction and false discovery rate control.
Selection bias arises when the data or results are chosen in a way that distorts inference. This may happen if only favorable analyses are reported or if hypotheses are formulated after inspecting the data. Both issues can make evidence appear stronger than it truly is.
8 Extensions and related concepts
8.1 Bayesian hypothesis testing
Bayesian hypothesis testing evaluates hypotheses using prior information and the observed data. Rather than focusing only on tail probabilities under a null model, it compares the relative support for competing hypotheses. Bayes factors and posterior probabilities are common tools in this framework.
This approach treats uncertainty in a different way from classical significance testing. It can incorporate prior knowledge explicitly and often produces conclusions that are directly interpretable as updated beliefs. However, results depend on the chosen prior and model specification.
8.2 Likelihood-based approaches
Likelihood-based methods compare hypotheses by examining how well they explain the observed data. The likelihood function measures the probability of the data as a function of the parameters. Hypotheses can then be assessed through likelihood ratios or related criteria.
These methods are closely tied to model fitting and estimation. They are widely used because they provide a coherent framework for comparing competing explanations. Likelihood-based reasoning also underlies many standard hypothesis tests.
8.3 Equivalence and non-inferiority hypotheses
Equivalence and non-inferiority hypotheses are used when the goal is to show that a difference is small enough to be acceptable. In equivalence testing, the aim is to demonstrate that two quantities are practically indistinguishable within a defined margin. In non-inferiority testing, the goal is to show that one method is not worse than another by more than an allowed amount.
These approaches reverse the usual testing logic. Instead of searching for evidence of difference, they seek evidence that any difference is negligible or limited. They are common in medicine, engineering, and other applied fields where practical thresholds matter.
8.4 Sequential testing and adaptive designs
Sequential testing allows data to be evaluated as they are collected rather than only at the end of a study. Adaptive designs may modify aspects of the study plan based on interim results while preserving statistical validity. These methods can improve efficiency and reduce unnecessary data collection.
Because repeated looks at the data can inflate error rates, special procedures are required. Stopping rules, monitoring boundaries, and adjusted significance thresholds are often used. When carefully designed, sequential approaches can maintain rigorous inference while adding flexibility.