1 Definition and basic concepts
Copy number variation is a form of structural genetic variation in which a stretch of DNA occurs in variable numbers among individuals. The affected segment may be absent, present once, or present multiple times, depending on the chromosome and the organism. CNVs are a major source of genomic diversity and may alter biological function by changing the amount of DNA available at a locus.
Unlike single-nucleotide changes, CNVs often span many bases and can include one or more genes, regulatory elements, or intergenic regions. Some are inherited stably through families, while others arise newly in a single generation. Their effects range from no obvious consequence to marked changes in development, physiology, or disease risk.
1.1 Structural variation in the genome
Structural variation refers to large-scale alterations in chromosome architecture, including deletions, duplications, insertions, inversions, and more complex rearrangements. Copy number variation is one major class within this broader category. Because CNVs change the number of DNA copies rather than only the sequence of a single base, they can have broader effects on gene content and genome organization.
1.2 Copy number changes
Copy number changes involve the gain or loss of DNA segments relative to a reference genome or to another individual. These changes may encompass entire genes, parts of genes, repeated elements, or noncoding intervals. The biological outcome depends on the size of the change, its genomic location, and whether it disrupts important functional elements.
1.2.1 Deletions
A deletion removes a genomic segment, reducing the copy number at that locus. Deletions can range from small losses affecting a few bases to large events spanning multiple genes. When a deleted region contains an essential gene or regulatory sequence, the result may be a strong phenotypic effect.
1.2.2 Duplications
A duplication produces an extra copy of a DNA segment. Duplicated material may lie adjacent to the original sequence or at another genomic position. Duplications can increase gene dosage, create novel gene arrangements, or supply raw material for evolutionary change.
1.2.3 Multiallelic copy number variation
Some loci have more than two common copy states in a population and are described as multiallelic CNVs. These regions may show a wide range of copy numbers, often because they contain repeated DNA prone to unequal recombination or replication slippage. Such variation is especially common in gene families and tandem repeat regions.
1.3 Size range and genomic scope
CNVs typically extend from a few hundred base pairs to several megabases. Small events may involve a single exon or a short repeat block, whereas large events can cover many genes and extensive regulatory landscapes. The genomic scope of a CNV strongly influences how it is detected and interpreted.
2 Formation and mechanisms
CNVs arise through several molecular mechanisms that alter chromosome structure during DNA replication or repair. Many events are shaped by local sequence features such as repeats, low-complexity tracts, or fragile genomic regions. The same mechanism may generate both benign and pathogenic variants, depending on where the change occurs.
2.1 DNA replication errors
Replication errors can produce copy number changes when the DNA polymerase machinery slips, stalls, or restarts incorrectly. Misalignment of repeated sequences may cause the newly synthesized strand to gain or lose material. These events are especially likely in regions containing repeated motifs or unstable sequence architecture.
2.2 Non-allelic homologous recombination
Non-allelic homologous recombination occurs when recombination takes place between similar sequences at different genomic positions rather than between matching alleles. This process can lead to deletions, duplications, or inversions, particularly in regions flanked by segmental duplications. It is a common source of recurrent CNVs with similar breakpoints across unrelated individuals.
2.3 Non-homologous end joining
Non-homologous end joining is a DNA double-strand break repair pathway that can reconnect broken DNA ends without extensive sequence homology. Because the repair process is error-prone, small insertions or deletions, and sometimes larger rearrangements, can result. Improper repair of multiple breaks may contribute to complex CNV formation.
2.4 Replication-based mechanisms
Replication-based mechanisms generate CNVs when the replication fork encounters obstacles and resumes synthesis using an abnormal template. These processes can create complex rearrangements with unusual breakpoint patterns. They are especially relevant in regions under replication stress.
2.4.1 Fork stalling and template switching
Fork stalling and template switching describes a model in which a stalled replication fork jumps to a nearby template and then returns to the original site. This can produce duplications, deletions, or more elaborate rearrangements. The resulting structure often reflects the sequence context surrounding the stalled fork.
2.4.2 Microhomology-mediated repair
Microhomology-mediated repair uses short stretches of matching sequence to join broken DNA ends. Because only limited homology is required, the process can generate rearrangements at diverse breakpoints. CNVs formed by this pathway may show brief regions of shared sequence at junctions.
3 Types of copy number variation
CNVs can be classified by structure, complexity, and relationship to nearby sequence. Different categories are often described to capture the way the DNA segment is arranged in the genome. This classification helps in interpretation, especially in medical genetics and genome analysis.
3.1 Simple CNVs
Simple CNVs involve a single gain or loss of a contiguous DNA segment. They are the most straightforward to describe and often have clear breakpoints. A simple CNV may delete one region or duplicate one region without additional rearrangement.
3.2 Complex CNVs
Complex CNVs contain more than one rearrangement event or a mixed pattern of gains and losses. These variants may involve multiple breakpoints, inserted fragments, or interwoven sequence segments. Their structure can be difficult to resolve with low-resolution methods.
3.3 Tandem duplications
Tandem duplications place an extra copy of a DNA segment directly adjacent to the original copy. The repeated copies are arranged in the same genomic neighborhood and may appear in head-to-tail orientation. Such duplications can alter expression levels or create repeated gene copies.
3.4 Segmental deletions
Segmental deletions remove a discrete chromosomal region, often containing one or more genes. The size may be small enough to affect only part of a gene or large enough to eliminate many neighboring elements. Segmental deletions are frequently discussed in clinical genetics because of their potential to disrupt dosage-sensitive loci.
3.5 Variable number tandem repeats
Variable number tandem repeats are repeated sequence units arranged in series, with the number of repeat copies differing among individuals. They are a classic example of copy number variation at a smaller scale. Because repeat number can fluctuate, these loci are useful in genetic analysis and population studies.
4 Detection and measurement
CNVs are identified using cytogenetic, molecular, and genomic approaches. The choice of method depends on the expected size of the variant, the purpose of the study, and the resolution required. Accurate measurement often benefits from combining more than one technique.
4.1 Cytogenetic methods
Cytogenetic methods examine chromosomes directly, usually at the level of visible structure or fluorescent signal. They are useful for large CNVs and for detecting rearrangements in clinical samples. However, they generally lack the resolution needed for smaller events.
4.1.1 Karyotyping
Karyotyping visualizes chromosomes under a microscope and can reveal large deletions, duplications, or other major structural changes. It is a traditional method in genetics laboratories. Its resolution is limited, so smaller CNVs usually escape detection.
4.1.2 Fluorescence in situ hybridization
Fluorescence in situ hybridization uses labeled DNA probes to bind specific genomic regions on chromosomes or within nuclei. Loss or gain of a probe signal can indicate a CNV at the target locus. This method is often used to confirm suspected abnormalities.
4.2 Molecular methods
Molecular methods estimate copy number by measuring DNA abundance at selected loci. They are often faster and more precise for targeted analysis than chromosome-based techniques. Such assays are valuable for validation and clinical testing.
4.2.1 Quantitative PCR
Quantitative PCR measures the relative amount of a DNA segment by comparing amplification signals with a reference locus. Differences in cycle thresholds can indicate copy number changes. The method is targeted and sensitive but limited to one or a few regions at a time.
4.2.2 Multiplex ligation-dependent probe amplification
Multiplex ligation-dependent probe amplification uses probe pairs that hybridize to target sequences and are then ligated and amplified. Relative signal intensity reflects copy number at many loci in a single assay. It is commonly applied to targeted diagnostic questions.
4.3 Genomic technologies
Genome-wide technologies can survey many loci simultaneously and are widely used in modern CNV discovery. They provide higher resolution than classical cytogenetics and can detect previously unrecognized variants. Interpretation still requires careful assessment of quality and context.
4.3.1 Array comparative genomic hybridization
Array comparative genomic hybridization compares test and reference DNA hybridization intensity across thousands of genomic probes. Gains and losses appear as changes in signal ratio. The method is effective for detecting unbalanced copy number changes across the genome.
4.3.2 SNP arrays
SNP arrays genotype many single-nucleotide markers and can infer CNVs from signal intensity and allelic patterns. They are especially useful in large-scale studies and clinical screening. Their performance is strongest for variants in well-covered regions.
4.3.3 Whole-genome sequencing
Whole-genome sequencing can detect CNVs by analyzing read depth, paired-end mapping, split reads, and assembly patterns. It offers broad coverage and can resolve breakpoints more precisely than many array-based approaches. Data analysis is computationally demanding and requires robust quality control.
4.4 Bioinformatic analysis
Bioinformatic analysis integrates raw signal or sequence data to identify candidate CNVs and estimate their boundaries. Algorithms consider coverage depth, discordant read pairs, and local sequence context. Downstream interpretation includes filtering technical artifacts, comparing to reference datasets, and evaluating likely biological impact.
5 Functional consequences
CNVs can influence function by changing how much product a gene makes, by disrupting sequence elements, or by altering chromosomal context. The same class of variant may have different outcomes in different tissues or genetic backgrounds. Functional interpretation therefore depends on both the variant and the surrounding genome.
5.1 Gene dosage effects
Gene dosage effects occur when changing the number of gene copies changes the level of RNA or protein produced. Increased copies may raise expression, while deletions may reduce it. Some genes tolerate dosage changes well, whereas others are highly sensitive.
5.2 Dosage-sensitive genes
Dosage-sensitive genes are genes for which copy number changes produce measurable consequences. These genes often participate in developmental pathways, transcriptional regulation, or other tightly controlled processes. Even a small imbalance can affect normal function.
5.3 Regulatory region disruption
CNVs may remove or duplicate enhancers, promoters, silencers, or other regulatory elements. In some cases, the coding sequence remains intact but expression changes because control regions are altered. Such effects may be tissue-specific and difficult to predict from sequence alone.
5.4 Effects on phenotype
The phenotypic impact of a CNV ranges from none to severe. Some variants subtly influence height, metabolism, or behavior, while others contribute to recognizable syndromes. The same CNV can also have variable expressivity among carriers.
5.4.1 Developmental variation
Developmental variation may arise when CNVs affect genes involved in embryonic growth or differentiation. Consequences can include differences in anatomy, growth rate, or neurodevelopment. The outcome often depends on the timing and dosage of gene expression.
5.4.2 Disease susceptibility
Certain CNVs increase susceptibility to disease without guaranteeing that disease will occur. They may raise risk by shifting gene dosage, altering immune pathways, or changing neural function. Such variants are often described as risk factors rather than direct causes.
5.4.3 Neutral variation
Many CNVs appear to have little or no detectable effect. They may occur in nonessential regions or in contexts where dosage changes are well tolerated. Neutral variants still contribute to genetic diversity and are useful in population studies.
6 Clinical significance
CNVs are important in medical genetics because they can underlie syndromes, contribute to diagnosis, or be incidental findings with no known effect. Clinical interpretation depends on size, gene content, inheritance, and correlation with phenotype. Classification typically distinguishes harmful, uncertain, and benign changes.
6.1 Pathogenic CNVs
Pathogenic CNVs are those with strong evidence for causing disease or contributing substantially to a disorder. They often disrupt critical genes, dosage-sensitive regions, or chromosome segments known to be associated with clinical syndromes. Identification of a pathogenic CNV can support diagnosis and guide management.
6.2 Benign and likely benign CNVs
Benign and likely benign CNVs are variants considered unlikely to cause disease. They may be common in healthy individuals or occur in genomic regions with low functional constraint. Recognition of benign variation helps avoid misinterpretation in diagnostic testing.
6.3 Copy number polymorphisms
Copy number polymorphisms are common CNVs found at appreciable frequency in populations. They are part of normal human genetic variation and often show limited clinical impact. These polymorphisms are useful markers in studies of diversity and inheritance.
6.4 Genomic disorders
Genomic disorders are conditions caused by recurrent structural changes, often including CNVs. They frequently involve regions prone to rearrangement because of repeated sequences or unstable chromosomal architecture. Many genomic disorders are defined by characteristic deletion or duplication intervals.
6.4.1 Microdeletion syndromes
Microdeletion syndromes result from loss of a small chromosomal segment that removes one or more important genes. Although the deleted region may be too small to see on a routine karyotype, it can have significant effects. Clinical features depend on the genes involved.
6.4.2 Microduplication syndromes
Microduplication syndromes arise when a small segment is present in extra copy number. The duplicated region may cause altered dosage or interfere with normal regulation. As with microdeletions, the phenotype varies with the exact genomic interval.
7 Population genetics and evolution
CNVs are shaped by population processes such as mutation, drift, and selection. Their frequencies can vary widely between populations and between species. Because they can alter traits quickly, CNVs are relevant to both short-term variation and long-term evolutionary change.
7.1 Allele frequency in populations
The frequency of a CNV in a population reflects its mutation rate, inheritance pattern, and selective effects. Rare CNVs may represent recent events, whereas common ones may have persisted over many generations. Population frequency data are important for distinguishing unusual variants from ordinary polymorphism.
7.2 Selection and drift
Natural selection can increase or decrease CNV frequency depending on fitness effects, while genetic drift can shift frequencies by chance. Some CNVs are neutral and spread mainly through drift, whereas others are maintained by balancing selection or removed by purifying selection. The strength of these forces varies by locus.
7.3 Species differences
CNV patterns often differ among species because of differences in genome structure, life history, and evolutionary history. Comparative studies can reveal lineage-specific expansions or losses of genes and repeated elements. These differences help explain variation in traits across organisms.
7.4 Adaptive evolution
CNVs can contribute to adaptation by changing gene dosage or by duplicating genes that later evolve new functions. In changing environments, an extra copy of a gene may confer an advantage, especially if it increases tolerance or metabolic flexibility. Over time, duplication can provide material for functional diversification.
8 Databases and reference resources
Reference resources catalog CNVs and related genome annotations to support research and clinical interpretation. They help users compare a variant against known examples and evaluate whether it has been previously reported. Consistent curation is essential because CNV interpretation depends heavily on context.
8.1 CNV databases
CNV databases collect reported copy number variants along with genomic coordinates, observed frequencies, and associated annotations. They are used to compare new findings with established variation. Such databases may include clinical reports, population datasets, or both.
8.2 Genome annotation resources
Genome annotation resources provide information about genes, transcripts, regulatory features, and repetitive regions. These resources are important for judging whether a CNV overlaps functional elements. Accurate annotation improves interpretation of potential biological effects.
8.3 Standards for interpretation
Standards for interpretation define how CNVs should be classified and reported. They typically incorporate variant size, gene content, inheritance, population frequency, and evidence of disease association. Standardization reduces inconsistency across laboratories and studies.
9 Research applications
CNVs are widely used in research to investigate genetic variation, gene function, and disease mechanisms. They are also valuable in comparative studies across populations and species. Because they can influence traits through dosage and structure, CNVs provide a direct link between genome architecture and phenotype.
9.1 Disease association studies
Disease association studies examine whether particular CNVs occur more often in affected individuals than in controls. These studies can reveal risk loci and highlight genes involved in development or physiology. Careful design is needed to distinguish causal effects from background variation.
9.2 Pharmacogenomics
Pharmacogenomics studies how genetic variation affects drug response, including metabolism, efficacy, and adverse effects. CNVs in drug-metabolizing enzymes or transporters can alter gene dosage and change how a person processes medications. This makes copy number assessment relevant to individualized treatment strategies.
9.3 Functional genomics
Functional genomics uses CNVs to probe the roles of genes and regulatory regions. Introducing deletions or duplications can help determine whether a locus is dosage sensitive or how expression changes with copy number. These experiments are informative in cell lines, model organisms, and primary tissues.
9.4 Comparative genomics
Comparative genomics examines CNV patterns across different genomes to study evolutionary conservation and divergence. Shared and unique copy number changes can indicate species-specific adaptations or conserved functional constraints. Such analyses also help identify genomic regions prone to rearrangement.