1 Sampling fundamentals
1.1 Population, target population, and sampling frame
In survey sampling, a *population* is the set of all units about which conclusions are desired. Often the *target population* is more specific, defined by geography, time period, and eligibility rules (for example, residents of a region during a given year). Because data collection cannot reach every unit with perfect completeness, practitioners rely on a *sampling frame*, a list or operational mechanism intended to represent the target population. Examples include household address lists, registries, or phone-number databases.
A key practical distinction is that a frame may differ from the target population. Those differences drive coverage error, a central quality issue in sampling. Good frame construction aims to minimize omissions and duplications while allowing efficient selection and contact procedures.
1.2 Units of analysis and sampling units
Survey results may concern different *units of analysis*—such as individuals, households, or establishments—depending on the research question. The *sampling unit* is the object selected in the sample design. A common pattern is multi-level selection: for instance, sample households first, then select one person within each household for interviewing. These stages affect how inclusion probabilities, weights, and variance calculations are defined.
Clarifying the relationship between sampling units and the eventual analysis unit helps avoid mismatches between design assumptions and statistical estimation procedures.
1.3 Sampling goals and types of estimates
Sampling goals determine the nature of estimates. Surveys may estimate population totals, means, proportions, medians (often via specialized methods), or distributions across categories. The target estimand can be *univariate* (one outcome) or *multivariate* (several related outcomes). Some designs prioritize accurate estimates for subgroups, while others focus on overall population averages.
In addition to point estimates, sampling is also used to quantify uncertainty through standard errors and confidence intervals. The practical target is not only unbiasedness in theory, but accuracy and precision under real-world constraints.
1.4 Key sources of error in surveys
Survey error is often described as the sum of multiple components, each arising from different stages of the process. While exact decomposition depends on the study, four recurring categories are coverage, selection, measurement, and nonresponse errors.
These errors may interact: for example, an omitted subgroup can also differ in willingness to respond, compounding bias.
1.4.1 Coverage errors
*Coverage error* occurs when the sampling frame fails to perfectly represent the target population. Common mechanisms include incomplete lists, outdated addresses, restricted registries, or overinclusive frames that contain ineligible units. Coverage error can produce biased estimates if omitted or duplicated groups have systematically different outcomes.
Coverage errors are typically addressed through improved frame updates, careful definition of eligibility criteria, and post-collection adjustments when auxiliary data are available.
1.4.2 Selection bias
*Selection bias* arises when the sample selection mechanism does not produce representative inclusion of all units, even after accounting for known sampling probabilities. In probabilistic designs, selection bias is usually conceptualized through correct weighting and inclusion probability modeling. In nonprobability designs, selection processes can correlate with outcomes and may not be fully recoverable without strong assumptions.
Bias can also appear when field implementation deviates from the intended design, such as applying substitution rules incorrectly.
1.4.3 Measurement error
*Measurement error* reflects imperfections in how responses are obtained and recorded. It can stem from question wording, recall limitations, interviewer effects, mode effects (online vs telephone), and coding errors. Measurement error affects both bias and variance, and it is typically addressed through questionnaire design, training, pretesting, and quality control.
Unlike coverage and selection errors, measurement error often requires substantive knowledge about the survey topic to mitigate effectively.
1.4.4 Nonresponse error
*Nonresponse error* occurs when sampled units do not provide usable data. When nonresponse differs across outcome-relevant characteristics, estimates can shift. Nonresponse includes both total refusal and inability to contact, and it may occur at item level (some questions unanswered) as well as at unit level (no interview).
Nonresponse error is managed through follow-up protocols, weighting adjustments, and, when appropriate, imputation techniques that support statistical inference.
2 Sampling designs
2.1 Probability sampling
*Probability sampling* means each unit has a known, nonzero chance of selection. This property supports design-based inference using inclusion probabilities and yields a clear framework for variance estimation. Probability designs also help clarify how weighting should be applied so that estimates reflect the target population.
Practical trade-offs include the complexity of implementation and variance inflation from certain clustering patterns.
2.1.1 Simple random sampling
*Simple random sampling* selects units uniformly from the frame so that each unit has the same probability of inclusion. It is conceptually straightforward and often used as a baseline. However, for large geographically dispersed populations, it can be logistically inefficient because it may require visiting many distant locations.
In practice, simple random sampling is most feasible for smaller frames or when units are naturally centralized.
2.1.2 Systematic sampling
*Systematic sampling* selects units by choosing a random start and then taking every k-th unit from an ordered frame. It reduces administrative burden compared with fully random selection. Properties depend on how the frame is ordered; if the ordering correlates with the outcome, the sample can become biased.
When ordering is essentially random relative to the target variable, systematic sampling can be nearly as efficient as simple random sampling.
2.1.3 Stratified sampling
*Stratified sampling* partitions the frame into strata that are believed to differ meaningfully on key characteristics, then samples within each stratum. This can improve precision because variability is reduced within strata. Strata can be defined using auxiliary data such as region, urbanicity, or demographic indicators.
Choice of allocation—how many units to select per stratum—strongly affects efficiency.
2.1.4 Cluster sampling
*Cluster sampling* selects groups (clusters) of units, then measures units within chosen clusters. Clusters may be geographic areas or organizational units. Because sampling happens in “bunches,” intracluster similarity can increase variance relative to fully dispersed samples.
Cluster designs are widely used because they greatly reduce travel and field costs.
2.1.5 Multistage sampling
*Multistage sampling* combines selection across multiple hierarchical stages, such as selecting regions, then households within selected regions, then individuals within households. This approach supports broad coverage with manageable operational costs. It also requires careful treatment of inclusion probabilities across stages.
Multistage designs are common in large-scale national surveys due to the need for operational scalability.
2.2 Nonprobability sampling
*Nonprobability sampling* does not guarantee known selection probabilities for each unit. Estimates may still be useful, particularly when the goal is exploratory, but inferential claims about the target population require caution. Nonprobability samples often rely on assumptions about the relationship between participation and the outcome.
Common benefits include lower cost and faster turnaround.
2.2.1 Convenience sampling
*Convenience sampling* selects units that are easy to reach, such as respondents available in a location or willing to participate through accessible channels. While efficient, it can strongly overrepresent certain groups and underrepresent others.
Results are typically best interpreted as describing the sample rather than representing the entire target population.
2.2.2 Purposive sampling
*Purposive sampling* selects units based on judgment about their relevance to the research question. This can be valuable for studying rare characteristics or ensuring coverage of particular types of cases. However, the selection mechanism is subjective and may introduce systematic distortions.
Statistical generalization is limited unless additional structure or auxiliary information is used.
2.2.3 Quota sampling
*Quota sampling* sets targets for subgroups (quotas) and recruits until those quotas are filled. Unlike stratified probability sampling, quota selection does not inherently define inclusion probabilities. If recruitment within quotas is not random, bias can remain.
Quota designs sometimes incorporate modeling or weighting to improve comparability, but assumptions about within-quota representativeness are central.
2.2.4 Volunteer and opt-in samples
*Volunteer and opt-in samples* recruit participants who choose to participate, often via online panels or advertising. The self-selection mechanism can correlate with attitudes or behaviors, making it challenging to extrapolate to the broader population.
Researchers sometimes correct for nonrepresentativeness using poststratification or calibration, but these adjustments require credible auxiliary information and careful evaluation.
2.3 Mixed and hybrid designs
Many modern surveys blend elements from different paradigms to meet operational realities.
2.3.1 Random + nonrandom components
Some studies use probability sampling for certain stages and nonprobability approaches for others, such as a probability-based frame but an opt-in component for additional modules. This can be seen as a hybrid of design-based and model-assisted thinking.
Inference must reflect the actual selection mechanism across stages, often combining weighting with modeling.
2.3.2 Web sampling frameworks
Web-based frameworks may combine random selection from address-based registers (probability-like) with web-only collection modes, or they may rely on opt-in recruitment (nonprobability). Key considerations include coverage of internet users and the responsiveness differences between recruited groups.
Web sampling also introduces mode effects that can alter measurement and response behavior.
3 Sample size and precision
3.1 Determinants of required sample size
Required sample size depends on the target precision, the variability of the outcome, and the sampling design. Higher outcome variability, more detailed subgroup analysis, and complex designs typically call for larger samples. Additionally, anticipated nonresponse affects the effective number of completed interviews.
Researchers often begin with a target margin of error or variance for a primary estimate, then inflate sample size to accommodate design complexity.
3.2 Estimation targets and margins of error
A *margin of error* summarizes the uncertainty around a point estimate under specified confidence levels. Sample size calculations require assumptions about the distribution of the estimator and, in practice, choices about conservative bounds. For proportions, common approaches treat unknown parameters using pilot estimates or worst-case approximations.
For means and totals, similar logic applies through standard error formulas or simulation-based planning.
3.3 Power and detectability considerations
In studies with hypothesis testing objectives, *power* quantifies the probability of detecting a difference of interest given a specified effect size and significance level. Detectability depends on the standard error, which itself is driven by the design and variance structure.
Even for descriptive surveys, power concepts can influence how large the sample must be for meaningful subgroup comparisons.
3.4 Design effect and effective sample size
Complex designs often yield higher variance than simple random sampling. The *design effect* expresses this inflation relative to an equal-sized simple random sample. When the design effect is greater than one, the *effective sample size* is smaller than the raw completed sample.
Design effect is influenced by clustering, stratification, weighting variability, and intraclass correlations.
3.5 Finite population considerations
When the sampling fraction is substantial, the *finite population correction* can reduce variance because sampling without replacement reduces uncertainty. In large populations with small sampling fractions, this correction is usually negligible.
Planning with finite population correction is particularly relevant in smaller or well-defined populations.
3.6 Allocation strategies
Allocation describes how the total sample is distributed across strata or other subcomponents. For stratified designs, allocation can be proportional (equal sampling fractions) or optimized using variance estimates, balancing cost and precision goals. For multistage designs, allocation across stages influences operational feasibility and statistical efficiency.
Good allocation reduces waste by assigning more sampling effort to parts of the population with higher variability or greater analytic importance.
4 Selection, weighting, and inference
4.1 Inclusion probabilities and weights
In probability sampling, each unit’s inclusion probability governs how it should contribute to estimates. The inverse of inclusion probability, under appropriate conditions, forms the basis of *design weights*. Weights ensure that units sampled with lower probabilities effectively represent larger portions of the population.
When designs are multistage or involve unequal probabilities, weights must be computed consistently with the design specification.
4.2 Weighting methods
Weighting aims to correct imbalance due to unequal selection and, when possible, address nonresponse and coverage issues using auxiliary information.
4.2.1 Design weights
*Design weights* arise directly from the sampling mechanism. They incorporate selection probabilities at each stage and adjust for unequal probabilities or oversampling. Using design weights preserves the intended population representation under correct implementation and response behavior that is missing at random with respect to design-related factors.
Design weights typically precede any additional adjustments for nonresponse or post-survey alignment.
4.2.2 Post-stratification
*Post-stratification* adjusts weights so that weighted sample totals match known population counts across post-defined groups. These groups may be derived from variables collected in the survey and known from external sources, such as census distributions.
Post-stratification can improve alignment, but it assumes that within-group representativeness is adequate for the outcome of interest.
4.2.3 Calibration weighting
*Calibration weighting* modifies weights so that they satisfy constraints related to auxiliary variables, often using a distance function to stay close to design weights. This framework generalizes post-stratification and can accommodate multiple constraints.
Well-chosen auxiliary variables can reduce bias and improve precision, though mis-specified constraints may not deliver the intended gains.
4.2.4 Nonresponse adjustment
*Nonresponse adjustment* increases weights for respondents within classes where nonresponse occurred, using a model or adjustment cells. Approaches range from response-rate weighting to more elaborate modeling-based adjustments.
The credibility of nonresponse adjustments depends on whether nonresponse is adequately captured by the adjustment variables and whether the adjustment model is stable.
4.3 Variance estimation
Variance estimation under complex designs differs from naive calculations that treat observations as independent and identically distributed.
4.3.1 Analytical (formula-based) variance
*Analytical variance estimators* use design information and assumptions to compute standard errors, often involving stratum and cluster structure. These methods can be fast but depend on correctness of modeling assumptions and availability of sufficient design detail.
They may be sensitive to small numbers of clusters or strata.
4.3.2 Replication methods
*Replication methods* estimate variance by repeatedly reanalyzing the data with systematically perturbed sampling weights or replicate samples. This includes approaches like jackknife-style delete-a-group and resampling schemes designed for complex surveys.
Replication methods are commonly used because they can flexibly handle complex weighting and estimation processes.
4.3.3 Bootstrap and jackknife approaches
*Bootstrap* and *jackknife* procedures adapt to survey designs by respecting clustering and stratification. The effectiveness depends on the implementation details and on whether resampling preserves the design dependence structure.
When applied correctly, these methods provide practical standard errors for a wide range of estimators.
4.4 Survey inference and confidence intervals
*Survey inference* converts estimates and standard errors into uncertainty statements, such as confidence intervals and hypothesis tests. Under complex sampling, inference should use variance estimates that reflect the design and weighting.
Confidence interval construction often follows approximate normality assumptions or uses transformations for bounded outcomes.
4.4.1 Ratio and regression estimators
Some survey estimands are estimated using *ratio* or *regression* estimators that use auxiliary totals. Ratio estimators scale an estimate using a related variable, often reducing variance when the auxiliary variable correlates with the outcome. Regression estimators model the outcome as a function of auxiliary variables, then use fitted relationships to improve precision.
These estimators can be effective, but correctness depends on stability of the relationships and on appropriate variance estimation.
4.4.2 Standard errors under complex designs
Standard errors under complex designs require accounting for clustering, stratification, weighting, and covariance among observations. Naive independence-based formulas can seriously underestimate uncertainty.
Modern survey software typically implements design-aware variance estimation techniques, but careful settings and design inputs are essential.
5 Practical implementation in field surveys
5.1 Building and maintaining sampling frames
Frame construction includes listing, deduplication, eligibility checks, and periodic updates. Because population composition changes over time, frames can become outdated, leading to both coverage loss and incorrect inclusion. Maintenance may rely on administrative records, commercial lists, or periodic census operations.
Good practice includes auditing frame quality, tracking known gaps, and aligning definitions between the frame and target population.
5.2 Randomization, tracking, and auditability
Sampling requires reproducible randomization procedures. Fieldwork systems should record selection parameters, contact attempts, call outcomes, and any deviations from the protocol. Auditability helps detect implementation errors and supports re-weighting or correction when necessary.
Documentation also allows future methodological comparisons and transparency in reporting.
5.3 Fieldwork operational constraints
Even with well-specified designs, real-world constraints influence data quality. Constraints include staffing, geographic logistics, interviewer workload, and time windows for contact.
Operational realities interact with sampling choices: for example, cluster sampling may reduce travel but increase variance and burden allocation complexity.
5.3.1 Substitution rules
Substitution rules specify what happens when a sampled unit is unavailable or ineligible. Rules may allow replacing a nonworking address with another nearby unit or selecting a fallback within the same cluster. Substitution can reduce costs but may introduce bias if substitutions are systematically different from originals.
To preserve interpretability, substitution procedures should be explicit and incorporated into the weighting or analysis plan when possible.
5.3.2 Call scheduling and contact strategies
Contact strategies determine the likelihood of reaching sampled units. Call scheduling affects responsiveness, particularly when different subgroups have different availability patterns. Strategies may include repeated attempts at varied times, multi-mode contacting (phone then web), and use of advance notifications when allowed.
Recording contact histories supports nonresponse adjustment and quality evaluation.
5.4 Handling nonresponse
Nonresponse handling includes both process adjustments during collection and post-collection statistical corrections.
5.4.1 Nonresponse follow-up
Nonresponse follow-up typically involves additional attempts, possibly using alternative contact modes or targeted outreach. Follow-up can be staged, prioritizing cases with higher expected response propensity or relevance to key estimates.
Careful follow-up balances diminishing returns with budget limits.
5.4.2 Imputation basics (as a bridge to inference)
*Imputation* fills missing values using information from observed responses and auxiliary variables. As a bridge to inference, imputation allows analysts to maintain sample size and reduce item nonresponse-related bias. Approaches include hot-deck methods, regression-based imputation, and multiple imputation frameworks that propagate uncertainty.
Imputation strategies require careful assumptions about missingness mechanisms and should align with the survey’s design and weighting structure.
5.5 Documentation and reproducibility
Reproducibility depends on documenting survey instruments, sampling design specifications, weight construction steps, and variance estimation methods. Version-controlled questionnaires, stable random seeds for selection where feasible, and archived processing scripts support auditing and reanalysis.
Clear reporting also helps external users judge whether assumptions and procedures were applied correctly.
6 Special topics in sampling
6.1 Sampling rare populations
Rare populations pose challenges because standard sampling strategies may yield too few observations for stable estimates. Outcomes may also be highly correlated with the rarity-defining characteristic.
6.1.1 Oversampling strategies
*Oversampling* selects a higher fraction of units likely to belong to the rare subgroup, improving data availability for subgroup analysis. This requires design weights or appropriate adjustments during estimation to retain target population representativeness.
Oversampling often uses frame variables correlated with rarity, such as location or administrative indicators.
6.1.2 Methods for small-area estimation
*Small-area estimation* targets reliable estimates for domains with limited sample sizes, such as local districts. Methods may combine model-based information with survey data, borrowing strength across areas. Approaches often rely on linking outcomes to auxiliary predictors and quantifying uncertainty across domains.
Because model assumptions matter, reporting model diagnostics and sensitivity analyses is common.
6.2 Overspecification and overcoverage effects
Overspecification occurs when sampling and estimation plans include overly detailed subgroup definitions relative to available sample size, which can inflate variance and complicate weighting. Overcoverage in the design or frame can also dilute efficiency by including many irrelevant units.
Mitigating these issues involves aligning stratification and domains with analytic needs and ensuring that frame quality supports the intended segmentation.
6.3 Missing data related to sampling
Missingness can arise from the sampling process itself and from the response process. Distinguishing these sources helps avoid inappropriate corrections.
6.3.1 Distinguishing sampling-based vs item nonresponse
*Sampling-based nonresponse* refers to missing entire units (no interview), while *item nonresponse* refers to unanswered questions among interviewed units. These categories can have different mechanisms and thus require different handling. Unit nonresponse affects inclusion and weighting, whereas item nonresponse affects estimator construction and imputation.
Clear metadata on which variables are missing and why supports appropriate modeling choices.
6.4 Ethical considerations in sampling practice
Ethical issues influence how sampling is conducted, especially regarding participation, privacy, and consent.
6.4.1 Informed consent and transparency (process-focused)
Informed consent involves informing participants about the purpose of the survey, what participation entails, and how data will be used. Transparency in sampling logistics includes clarifying how contact is initiated and how eligibility is determined. Process-focused transparency aims to support voluntary participation without exerting undue pressure.
Ethical practice also includes handling refusals respectfully and documenting consent procedures.
6.4.2 Privacy-preserving sampling logistics
Privacy-preserving logistics address how contact information and selection processes are managed. Measures can include secure storage of frame data, limited access controls, separation of identifying information from survey responses, and anonymization in analytic datasets.
When web sampling is used, researchers also consider data retention, secure authentication practices, and minimization of personally identifiable information.