1 Selective sweep basics

1.1 Definitions and key concepts

A selective sweep is an evolutionary process in which a beneficial allele increases in frequency because individuals carrying it have higher fitness than those without it. As selection drives the allele upward, linked genetic variation in the surrounding region can be carried to higher frequencies as well, often causing a temporary reduction in genetic diversity near the selected site.

The key idea is that genomes are inherited as largely intact blocks over short timescales, but recombination gradually breaks these associations. Selection tends to remove variation linked to the favored allele, while recombination works against this by creating new haplotypes.

1.2 Genetic hitchhiking and linkage

Genetic hitchhiking refers to the phenomenon where neutral or weakly selected variants rise in frequency not because they are beneficial, but because they are genetically linked to the selected allele. Linkage is mediated by physical proximity on chromosomes: loci close together tend to be inherited together, making them more likely to hitchhike during the sweep.

The rate at which recombination separates the selected allele from neighboring variants determines how strongly and how far the “hitchhiking” effect extends along the chromosome.

1.3 Role of allele frequency dynamics

Selective sweeps are characterized by allele frequency change over time. In many models, the rise of the advantageous allele is approximately deterministic when it is sufficiently common, but early spread can be stochastic when the allele is rare. The resulting time course influences how much genetic diversity is removed and how the signatures appear in present-day samples.

Whether the allele begins at high frequency, emerges as a new mutation, or experiences partial fixation alters the dynamics and the observable patterns across the genome.

1.4 Relationship to genetic diversity

Genetic diversity around a selected locus is expected to decline during a sweep because lineages carrying alternative haplotypes are less likely to persist. The strength of diversity reduction depends on selection intensity, the time elapsed since the sweep, recombination, and population history. In addition, the sweep can shift allele frequency spectra, producing an excess of certain intermediate-frequency variants and an altered balance of rare and common alleles near the target.

Demographic events such as bottlenecks can also reshape diversity, so selective-sweep interpretation typically requires jointly considering selection and demography.

2 Types of selective sweeps

2.1 Hard sweeps

A hard sweep describes a case where adaptation proceeds largely from a single origin of the beneficial allele, such as a new mutation that increases rapidly in frequency. Because all sampled copies of the allele trace back to a near-identical starting haplotype, the surrounding region often shows a strong, localized signature: long haplotypes with reduced diversity and high frequencies.

2.1.1 Origin from new mutations

When a beneficial allele arises via a new mutation and is sufficiently advantageous, selection can cause it to rise before alternative copies appear. The resulting trajectory creates a characteristic pattern of reduced heterozygosity and elevated haplotype similarity around the selected site.

This scenario is often easier to interpret because the sweep tends to “collapse” variation around the beneficial allele toward a dominant genomic background.

2.1.2 Origin from standing variation

In standing variation, the beneficial allele exists in the population at appreciable frequency before selection becomes strong. If the allele’s fitness increases due to environmental change, selection can drive it upward from multiple genetic backgrounds already present, though in some cases the trajectory still resembles a relatively sharp sweep if the allele’s initial distribution is limited.

Signatures can be weaker than those from strict single-origin sweeps because the allele is not necessarily associated with one unique haplotype.

2.2 Soft sweeps

A soft sweep involves an increase in frequency of beneficial alleles from more than one genetic origin. Instead of one dominant haplotype expanding, multiple haplotypes carrying the advantageous variant rise concurrently or sequentially.

Soft sweeps are often expected when the beneficial allele is old (standing variation) or when the mutation rate and population size allow multiple independent beneficial origins.

2.2.1 Multiple independent origins

If several independent mutations generate the beneficial allele (or closely related alleles) around the same time, selection can increase their collective frequency. The resulting genomic region often retains more diversity than in a hard sweep because haplotypes are not all identical at the time of sampling.

This leads to different expectations for haplotype structure and linkage disequilibrium.

2.2.2 Selection on standing variation

Selection on standing variation can produce soft sweep patterns when the favored allele is embedded in multiple haplotypes prior to selection. Even if selection acts strongly, the presence of multiple backgrounds limits the extent of haplotype homogenization.

Consequently, markers near the selected site may show a moderate reduction in diversity rather than the pronounced signature typical of a hard sweep.

2.3 Partial sweeps and incomplete fixation

A partial sweep occurs when selection does not lead to near-fixation in the sampled time window. This can happen when the allele is beneficial but not strongly enough to fix, when the environment changes again, or when sampling occurs soon after the allele starts increasing.

Incomplete fixation can produce signatures that resemble weaker sweeps overall, with allele-frequency distributions showing intermediate frequencies and haplotype-based statistics reflecting ongoing, not completed, selection.

3 Mechanistic drivers

3.1 Selection coefficients and fitness landscapes

The selection coefficient quantifies the relative fitness advantage of the beneficial allele. Larger values typically yield faster rises in frequency, creating stronger hitchhiking effects and more distinctive reductions in diversity. However, the effective trajectory also depends on the fitness landscape, including whether there are dominance effects, epistasis, or multiple adaptive steps.

If selection is frequency-dependent or interacts with other loci, the sweep dynamics can deviate from simple monotonic increase.

3.2 Recombination and genomic neighborhood size

Recombination reshuffles associations between the selected allele and its neighbors. High recombination rates shorten the genomic neighborhood affected by hitchhiking, while low recombination allows longer haplotype blocks to be carried along, increasing the spatial extent of sweep signals.

Because recombination rates vary across the genome, the same selective event can look different depending on local recombination environments.

3.3 Population size and stochastic effects

Population size influences genetic drift and the probability that a new beneficial mutation escapes loss due to chance. In smaller populations, stochastic effects can dominate early dynamics, potentially producing variable sweep outcomes, including failure to establish or slower-than-expected frequency trajectories. Larger populations reduce drift relative to selection and can lead to more consistent sweep patterns, with higher opportunities for multiple origins.

The timing of sampling relative to the sweep also interacts with stochasticity: a sweep can be detectable only during certain phases.

3.4 Mutation supply and timing

The mutation supply determines how often potentially beneficial alleles appear and how frequently soft sweep scenarios arise. When multiple advantageous mutations are possible and arise over overlapping timescales, selection may increase variants from multiple origins, generating soft sweep signatures.

Timing matters because selection onset relative to allele age governs whether standing copies exist and whether recombination has already reshaped haplotypes before the sweep intensifies.

3.5 Demography as a confounder

Demographic history—such as population bottlenecks, expansions, migration, and structure—can mimic or obscure sweep signals. Events that reduce effective population size increase genetic drift, creating local changes in allele frequencies and haplotype structure even without selection. Conversely, expansions can affect the spectrum of rare variants.

Therefore, sweep inference typically requires demographic modeling or approaches robust to demographic uncertainty.

4 Observable genomic signatures

4.1 Changes in allele frequency spectra

Selective sweeps can alter the distribution of allele frequencies. Near a beneficial site, selection may increase the frequency of linked variants while reducing variation that would otherwise remain at intermediate or low frequencies. The resulting allele frequency spectrum can include an excess of specific frequency classes relative to neutral expectations.

These patterns are sensitive to sweep type and time since selection began.

4.2 Reduction of heterozygosity

A hallmark of selective sweeps is reduced heterozygosity around the selected locus due to the sweeping haplotype dragging linked lineages upward in frequency. The reduction is often localized and decays with distance from the target as recombination breaks associations.

The depth of heterozygosity reduction depends on how rapidly the allele increased, how long it has been segregating, and whether the sweep is hard or soft.

4.3 Linkage disequilibrium patterns

Linkage disequilibrium (LD) measures non-random associations among alleles at different loci. During a sweep, LD around the selected site can be elevated because recombination has not yet fully disrupted the association between the favored allele and neighboring variants. Patterns of LD decay with distance can therefore carry information about sweep timing and strength.

In soft sweeps, LD may appear weaker or more complex because multiple haplotype backgrounds can increase together.

4.4 Extended haplotype homozygosity

Extended haplotype homozygosity refers to unusually long stretches where sampled haplotypes are similar, reflecting recent common ancestry created by a sweep. Because recombination fragments these blocks over time, the length of homozygous segments can inform the approximate recency of selection.

Different statistics quantify the extent of these haplotypes, often emphasizing the contrast between long haplotypes near selected sites and shorter haplotypes elsewhere.

4.5 Site frequency and spatial patterns across the genome

Sweep effects can vary across the genome due to local recombination rate heterogeneity and varying functional densities. In addition, multiple sweeps can overlap, producing composite patterns that complicate interpretation.

Spatially, selected regions may show clustered signals—such as concordant shifts in several summary statistics—while distant regions should resemble neutral expectations except where other selective processes occurred.

5 Detection and inference methods

5.1 Summary-statistic approaches

Summary-statistic methods scan genomes for deviations from neutral expectations using statistics such as measures of heterozygosity, allele-frequency shifts, LD-based patterns, and haplotype length metrics. These approaches are often computationally efficient and can provide candidate regions for follow-up.

Their performance depends on assumptions about selection timing, population demography, recombination, and the strength of the sweep.

5.2 Haplotype-based methods

Haplotype-based methods explicitly use phased haplotype structure to detect regions where long haplotypes are unusually common. Because hard sweeps produce strong haplotype similarity, these methods can be particularly informative for recent selection.

For soft sweeps and partial sweeps, haplotype signals may be subtler, requiring statistics that can capture partial homogenization or multiple origins.

5.3 Composite likelihood and model-based frameworks

Composite likelihood approaches combine information across loci while allowing certain parameters—such as selection strength, recombination rate, and demographic history—to be estimated or integrated. Model-based frameworks simulate expected patterns under specific evolutionary scenarios and compare them to observed data.

These strategies can improve interpretability and quantify uncertainty, though they may require careful specification of assumptions.

5.4 Machine learning and classification pipelines

Machine learning methods treat sweep detection as a classification or regression problem. Models may use engineered features derived from summary statistics or use learned representations from genomic data. When trained on simulated datasets spanning a range of demographic histories and selection parameters, such pipelines can generalize to new data and potentially distinguish sweep types.

However, performance is contingent on the realism of the training simulations and the match between simulated and empirical linkage, sampling, and error structure.

5.5 Validation with simulations

Simulations are central for assessing detection power, calibrating false discovery rates, and examining sensitivity to confounding processes such as demographic change and recombination-rate mis-specification. By generating synthetic genomes under controlled selection and neutral scenarios, researchers can evaluate which signatures are most reliable for given conditions.

Validation typically includes parameter sweeps to understand performance across a broad space of selection strengths, allele ages, and population sizes.

6 Modeling selective sweeps

6.1 Forward-time vs coalescent simulations

Forward-time simulations track allele frequencies and haplotypes generation-by-generation, supporting complex models with selection, recombination, and varying demography. Coalescent-based simulations work backward from sampled genomes to ancestral processes, often enabling efficient inference under neutral and some selection approximations.

Choosing between these approaches depends on the complexity of the scenario, computational constraints, and the need to model detailed haplotype structure.

Many analytical models use Wright–Fisher assumptions, including discrete generations, random mating, and binomial sampling of alleles. Under these approximations, selection and recombination can be incorporated in mathematically tractable ways, yielding predictions for allele-frequency trajectories and patterns of linkage.

More elaborate models relax assumptions, such as allowing overlapping generations or structured populations.

6.3 Approaches for hard vs soft sweep regimes

Modeling differs between hard and soft regimes because the number of allele origins changes the structure of haplotypes. Hard sweep models often assume a single founding haplotype, leading to strong predictions about hitchhiking and homozygosity. Soft sweep models incorporate multiple initial copies or recurrent mutation, which yields less homogenization and more retained variation.

Inference frameworks often include parameters that interpolate between these regimes.

6.4 Incorporating recombination maps

Recombination maps provide spatially varying recombination rates, improving realism of predicted LD decay and haplotype lengths. Without such maps, models may misestimate the extent of linked-region effects, leading to biased selection parameter estimates or reduced detection power.

In practice, recombination maps are derived from empirical data and may vary between populations, so using an appropriate map is important for cross-population studies.

6.5 Handling demographic history in inference

Because demography can generate sweep-like signatures, many inference pipelines incorporate demographic parameters or use approaches that jointly fit selection and demographic models. Some methods use flexible demographic priors, while others rely on pre-estimated demographic models from neutral loci.

Robust inference often requires separating selection signals from demographic-driven changes in allele frequencies and haplotype structure.

7 Practical considerations in genomic studies

7.1 Sampling design and coverage biases

The detectability of sweep signatures depends on how individuals are sampled and how representative the sample is of the population. Unequal sampling across subpopulations, limited geographic coverage, or uneven time sampling (for ancient DNA) can distort inferred haplotype frequencies.

Sequencing and coverage depth can also affect the ability to call variants reliably, shaping downstream statistics.

7.2 Effects of sequencing and genotyping error

Errors in variant calling can reduce signal-to-noise ratios in LD and haplotype-based analyses. Systematic errors may create spurious correlations or mask real associations, particularly in low-frequency variants and genomic regions with complex mapping.

Quality control steps and error-aware modeling can mitigate these effects.

7.3 Window size and genome scan resolution

Genome scans often compute statistics in sliding windows, and the chosen window size influences sensitivity. Small windows may capture fine-scale signals but can be noisy, while large windows may dilute localized effects or blend signals from nearby events.

Resolution considerations also relate to linkage disequilibrium scales, which depend on recombination and local genomic architecture.

7.4 Multiple testing and false discovery control

Genome-wide scans test many hypotheses simultaneously, so raw p-values or scores require correction for multiple comparisons. Controlling false discovery rates or using permutation-based thresholds helps reduce the chance of false positives.

The correction method can influence which candidate regions survive, particularly when the distribution of test statistics is not uniform under the null.

7.5 Interpreting signals near recombination hotspots

Near recombination hotspots, LD can decay rapidly, shortening haplotype signatures and potentially weakening certain sweep metrics. However, hotspots can also create complex LD patterns that mimic partial recombination and may affect interpretations of sweep boundaries.

Accounting for local recombination rate variation is therefore important when assigning functional relevance to detected regions.

8 Case study patterns (generalized, non-specific)

8.1 Sweeps in rapidly changing environments

In environments that shift over relatively short evolutionary timescales, beneficial alleles can increase quickly, leaving stronger signatures that persist for measurable periods. Such conditions can increase the likelihood of detectable sweeps, especially when selection acts on standing variation that is already segregating.

General patterns often show clearer haplotype and diversity reductions near loci under selection, though demographic noise remains a concern.

8.2 Interpreting mixed or overlapping selective events

Real genomes can experience multiple selective episodes, sometimes affecting nearby loci or occurring within overlapping time windows. Overlapping sweeps can produce mixed signatures, such as heterozygosity reductions paired with allele frequency changes that do not match a single-sweep model.

Interpretation typically benefits from multi-locus modeling, haplotype decomposition, or conditional analyses that attempt to separate event histories.

8.3 Comparing sweep strength across loci

Sweep strength varies across genomic regions due to differences in selection intensity, recombination environment, target functional effects, and the local mutational context. When comparing loci, researchers often standardize expectations using local recombination rates and incorporate demographic background variation.

Comparisons can identify whether some genes or pathways show repeated adaptation signals, though such interpretations require careful statistical control.

8.4 Relating candidate regions to functional categories

Candidate sweep regions are frequently mapped to genes, regulatory elements, or conserved functional units using annotation resources. To connect signals to biology, researchers assess whether detected regions are enriched for particular functional categories or overlap with known regulatory motifs.

Functional interpretation is strengthened when genetic signals align with independent evidence such as expression changes, protein-coding impacts, or pathway-level coherence.