1 Definition and classification
1.1 General definition
A structural variant is a genomic change that alters the organization of DNA segments rather than replacing a single nucleotide with another. Such changes can involve stretches of DNA from a few dozen bases to many millions, and they may modify genes, regulatory sequences, or the physical arrangement of chromosomes. Structural variants are found in both inherited and newly arisen changes, and they form a major class of genome variation studied in genetics, molecular biology, and medicine.
1.2 Size-based distinctions
Structural variants are commonly distinguished from smaller sequence changes by their length and by the type of genomic alteration involved. Although no absolute boundary is used in every context, the term usually refers to rearrangements larger than typical single-base substitutions and short insertions or deletions. Some workflows also distinguish variants by whether they are detectable by sequence-based methods, cytogenetics, or array technologies.
1.2.1 Comparison with small variants
Small variants include single-nucleotide variants and short insertions or deletions that usually affect only a limited stretch of DNA. By contrast, structural variants can span multiple exons, whole genes, or broader chromosomal regions. Because of their scale, structural variants are more likely to change gene dosage, disrupt long-range regulation, or create novel junctions between distant genomic segments.
1.3 Major categories of structural variants
Structural variants are grouped according to the type of rearrangement they produce. The main categories include loss, gain, inversion, and exchange of DNA segments, along with more complex rearrangements that combine several events. In practice, the same variant may be described differently depending on the resolution of the assay and the reference genome used.
1.3.1 Deletions
A deletion removes a segment of DNA from the genome. Deletions may range from a small internal loss within a gene to a large chromosomal segment spanning many genes. Their effects depend on size and location, and they can eliminate coding regions, disturb splice sites, or remove regulatory elements.
1.3.2 Duplications
A duplication results when a DNA segment is copied and inserted again in the genome. Duplicated regions may lie adjacent to the original sequence or elsewhere in the genome. Duplications can increase gene dosage, alter expression levels, or provide material for evolutionary divergence.
1.3.3 Insertions
An insertion introduces additional DNA into a genomic location. The inserted sequence may originate from another part of the genome, from a mobile element, or from an external template such as viral DNA in some contexts. Insertions can disrupt genes if they occur within coding or regulatory regions.
1.3.4 Inversions
An inversion occurs when a DNA segment breaks at two points and reinserts in the opposite orientation. The total amount of DNA may remain unchanged, but the rearrangement can interrupt genes or alter regulation by changing the spatial context of nearby sequences.
1.3.5 Translocations
A translocation transfers a DNA segment to a different chromosomal location, often involving an exchange between nonhomologous chromosomes. Balanced translocations may preserve total DNA content, whereas unbalanced forms can lead to gains or losses of material. Translocations are especially important in cancer genetics and some inherited rearrangements.
1.3.6 Copy number variants
Copy number variants are gains or losses of DNA segments that change the number of copies present in the genome. They include many deletions and duplications identified at a scale large enough to influence dosage. The term is often used in human genetics when copy number change is the main measurable feature.
1.4 Complex structural rearrangements
Some genomic alterations do not fit neatly into a single class. Complex structural rearrangements may combine multiple deletions, duplications, inversions, and translocations in one event. These patterns can arise from clustered DNA breakage and repair, replication errors, or repeated cycles of rearrangement, and they often require high-resolution methods to resolve.
2 Molecular mechanisms
Structural variants arise through several molecular pathways that affect how DNA breaks, replicates, and recombines. Many mechanisms involve imperfect repair of damaged DNA, while others reflect template switching during replication or exchange between repetitive sequences. Mobile DNA elements also contribute to new structural changes.
2.1 DNA break formation and repair
Double-strand breaks are a common starting point for many structural variants. When broken DNA ends are repaired inaccurately, fragments may be lost, joined in the wrong order, or reattached to the wrong chromosome. The final structure depends on the repair pathway and the availability of matching DNA ends.
2.1.1 Non-homologous end joining
Non-homologous end joining repairs double-strand breaks by directly ligating broken DNA ends with little or no requirement for long sequence matching. This process is efficient but can be error-prone, especially when ends are damaged or multiple breaks occur near one another. Small insertions, deletions, and more complex rearrangements can result.
2.1.2 Homologous recombination
Homologous recombination uses a similar DNA template to repair breaks accurately. When recombination occurs between the wrong sequences or in an unusual genomic context, it can create rearrangements instead of faithful repair. Misaligned recombination is a major source of deletions, duplications, and some translocations.
2.2 Replication-based mechanisms
During DNA replication, the copying machinery may stall, detach, or switch templates. These replication-associated processes can generate structural changes without requiring a classic double-strand break as the initial event. They are especially relevant in regions of complex sequence architecture.
2.2.1 Fork stalling and template switching
Fork stalling and template switching occurs when a replication fork pauses and the polymerase transiently copies from an alternative template before returning to the original strand. This can produce duplications, deletions, and rearranged junctions. The mechanism is often associated with regions that are difficult to replicate.
2.2.2 Microhomology-mediated break-induced replication
Microhomology-mediated break-induced replication relies on very short matching sequences to initiate repair or restart replication. Because only minimal homology is needed, the process can join unrelated DNA segments and generate deletions or complex junctions. It is considered an error-prone pathway that can amplify structural instability.
2.3 Recombination between repetitive elements
Repetitive DNA creates many opportunities for mispairing because similar sequences occur in multiple genomic locations. Recombination between repeats can delete, duplicate, or invert the DNA between them. Such events are common in genomes rich in segmental duplications or other repeated elements.
2.3.1 Non-allelic homologous recombination
Non-allelic homologous recombination occurs when recombination takes place between similar but nonmatching repeat sequences. Because the repeats are not true alleles at the same locus, the exchange can shift the position of intervening DNA. This mechanism is a frequent cause of recurrent structural variants with similar breakpoints.
2.4 Mobile element activity
Mobile elements can move within the genome and occasionally carry nearby DNA with them or create insertional changes at new sites. Their insertion may disrupt coding regions, alter transcription, or generate local rearrangements. In some cases, mobile element activity also provides repeated sequences that later serve as substrates for recombination.
3 Genome and chromosome effects
Structural variants influence genome function through several routes. They may interrupt genes directly, change the number of gene copies, move sequences into new chromosomal neighborhoods, or alter larger-scale chromosome organization. The biological outcome depends on the size, position, and type of rearrangement.
3.1 Gene disruption
When a structural variant cuts through a gene or removes part of it, normal gene function may be lost or altered. Disruption can affect exons, introns, splice signals, promoters, or termination regions. Even if the coding region remains intact, changes in gene architecture may interfere with proper expression.
3.1.1 Frameshift and truncation consequences
If a deletion or insertion changes the reading frame, translation may proceed incorrectly, often producing a premature stop codon. Truncation can also occur when a rearrangement removes part of a coding region. Such outcomes frequently reduce protein function or trigger RNA surveillance pathways that eliminate abnormal transcripts.
3.2 Dosage changes
Structural variants can raise or lower the amount of gene product by changing how many copies of a gene are present. Dosage effects are particularly important for genes that require tightly balanced expression. Both loss and gain of copy number may have biological consequences.
3.2.1 Haploinsufficiency
Haploinsufficiency refers to a situation in which one functional copy of a gene is not enough for normal activity. A deletion that leaves only a single copy can therefore cause a phenotype even when the remaining copy is intact. This effect is common in dosage-sensitive genes.
3.2.2 Triplosensitivity
Triplosensitivity occurs when having three copies of a gene produces an abnormal effect. A duplication can then be harmful because excess gene product disturbs cellular balance. Triplosensitive genes are often involved in development, regulation, or dosage-sensitive pathways.
3.3 Position effects
A structural variant may leave a gene structurally intact but relocate it relative to enhancers, silencers, or chromatin domains. In such cases, the main consequence is not loss of sequence but altered genomic context. The result can be underexpression, overexpression, or inappropriate timing of gene activity.
3.3.1 Regulatory element displacement
When a rearrangement separates a gene from its normal regulatory elements, expression may change markedly. Conversely, a gene may be placed near a new enhancer or silencer that changes its activity. These displacement effects can be subtle or profound depending on the regulatory architecture.
3.4 Chromosomal architecture changes
Large rearrangements can disturb higher-order chromosome organization, including local folding patterns and long-range contacts. Such changes may influence how genes interact with enhancers or how regions of the genome are replicated and repaired. In extreme cases, the overall stability of a chromosome can be compromised.
4 Detection and analysis
Structural variants are detected with methods that examine chromosomes, DNA copy number, or sequence reads. No single technology identifies all classes equally well, so studies often combine multiple approaches. Analysis typically includes variant calling, filtering, annotation, and interpretation in the context of genomic function.
4.1 Cytogenetic methods
Cytogenetic techniques visualize chromosomes or large chromosomal segments directly. They are useful for large rearrangements and for cases where the overall structural context matters more than base-level detail. Their resolution is lower than that of sequencing, but they remain important in clinical and research settings.
4.1.1 Karyotyping
Karyotyping examines stained chromosomes under a microscope to detect large deletions, duplications, inversions, and translocations. It is especially effective for changes large enough to alter chromosome banding patterns. Because it surveys the whole set of chromosomes at once, it can reveal broad structural abnormalities.
4.1.2 Fluorescence in situ hybridization
Fluorescence in situ hybridization uses labeled DNA probes to detect specific genomic regions on chromosomes or in interphase nuclei. It can confirm suspected rearrangements, localize breakpoints approximately, and identify copy number changes at targeted loci. The method is often used when a focused question requires direct visualization.
4.2 Sequencing-based methods
Sequencing-based approaches identify structural variants by examining how DNA reads align to a reference genome. They can provide high resolution and breakpoint-level information, especially when read length and coverage are sufficient. These methods are now central to structural variant discovery.
4.2.1 Short-read sequencing
Short-read sequencing generates large numbers of relatively short DNA fragments that are aligned computationally. It is effective for many copy number changes and some breakpoint signatures, but repetitive regions and complex rearrangements can be difficult to resolve completely. Interpretation often depends on read depth and discordant alignment patterns.
4.2.2 Long-read sequencing
Long-read sequencing produces much longer contiguous reads, improving the detection of insertions, repeat-associated changes, and complex rearrangements. Because long reads can span breakpoints, they often reveal structural detail that short reads miss. The approach is especially valuable in repetitive or highly rearranged regions.
4.2.3 Paired-end and split-read analysis
Paired-end analysis compares the expected distance and orientation between two reads originating from the same DNA fragment. Split-read analysis identifies reads that align to more than one genomic location, indicating a breakpoint. Together, these signals provide strong evidence for structural variants and help refine breakpoint positions.
4.3 Array-based methods
Array technologies assess DNA copy number across the genome using hybridization or genotype intensity measures. They are particularly useful for gains and losses rather than balanced rearrangements. Although they do not usually define breakpoints precisely, they provide efficient genome-wide screening.
4.3.1 Comparative genomic hybridization
Comparative genomic hybridization compares test and reference DNA to detect copy number differences across genomic regions. Array-based versions offer improved resolution over earlier chromosome-based approaches. The method is widely used for identifying deletions and duplications.
4.3.2 SNP arrays
Single-nucleotide polymorphism arrays measure allele patterns and signal intensity across many loci. They can reveal copy number changes, regions of absence of heterozygosity, and some larger rearrangements. Their combined genotype and intensity information is useful in clinical genetics and population studies.
4.4 Computational calling and annotation
Bioinformatic analysis converts raw experimental data into candidate variants and then assigns biological meaning. Because structural variant signals can be noisy or ambiguous, computational steps are essential for accuracy. Annotation links detected changes to genes, regulatory regions, and known variant resources.
4.4.1 Variant filtering
Variant filtering removes low-confidence calls and probable artifacts using criteria such as read support, mapping quality, and consistency across samples. Filters may also exclude common technical artifacts associated with repetitive DNA or alignment errors. The goal is to retain variants with sufficient evidence for further study.
4.4.2 Interpretation pipelines
Interpretation pipelines combine detection, annotation, prioritization, and reporting into a structured workflow. They may evaluate predicted gene impact, overlap with known disease loci, inheritance pattern, and population frequency. In clinical settings, these pipelines help determine whether a structural variant is likely relevant to a patient’s phenotype.
5 Biological and medical significance
Structural variants are important because they contribute to normal genomic diversity and to many biological and medical phenotypes. Their effects range from neutral variation to severe disease, and they can influence gene expression, development, and evolutionary change. Some are inherited across generations, while others arise spontaneously in a person’s cells.
5.1 Contribution to genetic variation
Structural variants are a major source of variation within and between individuals. Because they can alter large DNA segments, they often have broader effects than single-base changes. This makes them significant contributors to genomic diversity and phenotypic differences.
5.1.1 Population diversity
Different populations may carry distinct structural variants that reflect ancestry, drift, and historical demographic events. Many of these variants are benign and may have little or no observable effect. Others can influence traits such as gene expression, metabolism, or susceptibility to specific conditions.
5.2 Role in inherited disorders
Some inherited disorders are caused by structural variants that disrupt genes or alter dosage. These changes may be transmitted in families or arise de novo in a gamete or early embryo. The resulting disorders often involve developmental, neurological, or congenital features.
5.2.1 Neurodevelopmental disorders
Structural variants can contribute to neurodevelopmental conditions by affecting genes involved in brain development, synaptic function, or neuronal signaling. Deletions and duplications are common findings in this area because dosage sensitivity is often important for neural processes. The clinical consequences may include developmental delay, learning differences, or other neurological features.
5.2.2 Developmental syndromes
Many developmental syndromes are associated with recurrent deletions or duplications of specific chromosomal regions. These rearrangements may disrupt multiple genes at once or alter dosage of a critical developmental regulator. The phenotypic outcome often depends on the exact genomic interval affected.
5.3 Role in cancer genomics
In cancer, structural variants may arise somatically within tumor cells and contribute to genomic instability. They can activate oncogenes, inactivate tumor suppressor genes, or create new gene combinations. Because they shape tumor biology, they are important in diagnosis, classification, and targeted research.
5.3.1 Somatic rearrangements
Somatic rearrangements are structural changes acquired after conception rather than inherited through the germline. They can accumulate during tumor evolution and create subclonal diversity. Their presence often reflects ongoing instability in the cancer genome.
5.3.2 Gene fusions
Gene fusions occur when rearrangements join portions of two different genes into one transcript or regulatory unit. The new combination may produce an abnormal protein or place a gene under control of a new promoter. Gene fusions are a well-known class of cancer-associated structural variants.
5.4 Evolutionary importance
Structural variants have played a substantial role in genome evolution by generating novelty in gene content, regulation, and chromosome organization. Over time, some rearrangements are lost, while others are retained because they are neutral or beneficial. Their accumulation contributes to lineage-specific differences among species.
5.4.1 Adaptation and speciation
Structural variants can support adaptation by changing gene dosage or modifying traits relevant to local environments. In some cases, rearrangements reduce recombination between populations and may contribute to reproductive isolation. Over long evolutionary periods, this can help distinguish species or subspecies.
6 Databases and nomenclature
Standardized naming and curated databases are essential for comparing structural variants across studies. Because these variants can be described at different levels of detail, common conventions help reduce ambiguity. Reference resources also support clinical review, population analysis, and research interpretation.
6.1 Variant nomenclature standards
Structural variant nomenclature aims to specify the type of change, its genomic coordinates, and the reference sequence used. Clear naming is especially important when the same variant may be represented differently by different platforms. Standardization improves communication between laboratories and databases.
6.1.1 Description conventions
Description conventions typically include chromosome location, breakpoints when known, and the nature of the rearrangement. Terms such as deletion, duplication, inversion, and translocation are used alongside coordinate-based notation. When precise breakpoints are unavailable, broader interval descriptions may be provided.
6.2 Reference databases
Reference databases collect structural variant records from population studies, clinical reports, and experimental investigations. They help determine whether a change is rare or common and whether it has been previously associated with a phenotype. Such resources are central to annotation and interpretation workflows.
6.2.1 Population variant catalogs
Population variant catalogs compile structural variants observed in groups of apparently healthy individuals or in large reference cohorts. These catalogs provide frequency information that helps distinguish common benign changes from rare candidate pathogenic variants. They are valuable for comparative studies and filtering.
6.2.2 Disease-associated variant resources
Disease-associated variant resources collect structural variants linked to inherited disorders, syndromes, or other medically relevant phenotypes. Entries may include genomic coordinates, reported effects, and supporting literature. These databases aid diagnosis and variant classification.
6.3 Reporting and clinical interpretation
Clinical reporting of structural variants requires careful assessment of gene content, predicted effect, inheritance, and supporting evidence. Interpretation often considers whether the variant matches the patient’s features and whether it is seen in unaffected populations. Because structural variants vary widely in size and mechanism, reports usually combine technical description with clinical significance.