1 Splice site definition and core elements

Splice sites are specific nucleotide sequences located at the boundaries between exons and introns in precursor messenger RNA (pre‑mRNA). They serve as recognition signals for the spliceosome, the large ribonucleoprotein complex that catalyzes intron removal and exon joining during RNA splicing. The three essential elements are the 5′ splice site (donor site), the 3′ splice site (acceptor site), and the branch point sequence. Accurate splice site recognition is critical for proper gene expression, and mutations in these sequences are a common cause of genetic disorders.

1.1 5′ splice site (donor site)

The 5′ splice site is located at the exon‑intron junction, immediately after the last nucleotide of the exon. In nuclear pre‑mRNA, the consensus sequence is typically AGGURAGU (where the vertical bar marks the cleavage site, R stands for purine, and the intron begins with GU). The GU dinucleotide at the intron start is nearly invariant in vertebrates. The 5′ splice site is recognized early in spliceosome assembly by the U1 small nuclear ribonucleoprotein (snRNP) through base‑pairing between the U1 snRNA and the pre‑mRNA sequence.

1.2 3′ splice site (acceptor site)

The 3′ splice site is located at the intron‑exon junction, immediately before the first nucleotide of the downstream exon. The consensus sequence is (Y)nNYAGG, where Y is a pyrimidine, n is a variable number (typically 10–20), N is any nucleotide, and the intron ends with AG. The terminal AG dinucleotide is highly conserved. The 3′ splice site is recognized by the U2 auxiliary factor (U2AF) and later by the spliceosome during the second catalytic step.

1.3 Branch point sequence and polypyrimidine tract

The branch point sequence is located within the intron, typically 18–40 nucleotides upstream of the 3′ splice site. In vertebrates, the consensus is YNYURAY (where Y is pyrimidine, N any nucleotide, R purine, A adenosine). The conserved adenosine (the branch point A) is the site of lariat formation during splicing. The polypyrimidine tract is a stretch of pyrimidines (mostly uracils and cytosines) between the branch point and the 3′ splice site; it facilitates binding of U2AF and is important for 3′ splice site recognition.

2 Splicing mechanism and spliceosome assembly

Splicing is a two‑step transesterification reaction catalyzed by the spliceosome. The spliceosome assembles stepwise on pre‑mRNA, beginning with recognition of splice sites by snRNPs and other auxiliary proteins. The assembly proceeds through a series of complexes: E (early), A, B, and C.

2.1 Spliceosome components (snRNPs U1, U2, U4/U6, U5)

The major spliceosome (for U2‑type introns) comprises five small nuclear ribonucleoproteins (snRNPs): U1, U2, U4, U5, and U6. Each snRNP contains a small nuclear RNA (snRNA) and associated proteins. U1 snRNA base‑pairs with the 5′ splice site; U2 snRNA base‑pairs with the branch point; U4 and U6 are extensively base‑paired and form a di‑snRNP with U5; U5 snRNA interacts with both exons. The catalytic core is formed by U6 and U2 snRNAs.

2.2 Stepwise assembly: E complex, A complex, B complex, C complex

Spliceosome assembly begins with formation of the E (early) complex, followed by the A complex (pre‑spliceosome), then the B complex (mature spliceosome), and finally the C complex (catalytic). Each step involves ATP‑dependent rearrangements and changes in RNA‑RNA interactions.

2.2.1 Recognition of 5′ splice site by U1 snRNP

In the E complex, U1 snRNP binds to the 5′ splice site via complementary base‑pairing between U1 snRNA and the pre‑mRNA. This binding is stabilized by SR proteins and other splicing factors. The recognition is initially reversible but becomes committed as assembly proceeds.

2.2.2 Recognition of branch point by U2 snRNP

In the A complex, U2 snRNP binds to the branch point sequence, displacing the branch‑point‑binding protein (BBP). U2 snRNA forms a short duplex with the branch point region, bulging out the conserved adenosine that will participate in the first catalytic step. This binding requires ATP and is assisted by U2AF, which binds the polypyrimidine tract.

2.2.3 Docking of tri‑snRNP and catalytic activation

The B complex forms when the pre‑assembled U4/U6·U5 tri‑snRNP docks onto the A complex, bringing the U5 and U6 snRNAs into proximity. Extensive conformational rearrangements then occur: U1 is released from the 5′ splice site, and U6 base‑pairs with U2 and the 5′ splice site. U4 is dissociated (U4/U6 unwinding), allowing U6 to form the catalytic center. This yields the activated B* complex, which then commits to the first catalytic step to form the C complex.

2.3 Two‑step transesterification reaction

Splicing proceeds through two sequential transesterification reactions, both catalyzed by the spliceosome’s RNA‑based active site (formed by U6 and U2 snRNAs). No external energy is consumed in the chemical steps; ATP is used only for assembly and remodeling.

2.3.1 First step: 5′ splice site cleavage and lariat formation

In the first step, the 2′‑hydroxyl group of the branch point adenosine attacks the phosphodiester bond at the 5′ splice site. This cleaves the pre‑mRNA, freeing the 5′ end of the intron as a linear molecule. The 5′ end of the intron becomes covalently linked to the branch point adenosine via a 2′‑5′ phosphodiester bond, forming a lariat structure. The upstream exon now has a free 3′‑hydroxyl group.

2.3.2 Second step: 3′ splice site cleavage and exon ligation

In the second step, the 3′‑hydroxyl group of the upstream exon attacks the phosphodiester bond at the 3′ splice site. This cleaves the intron at its 3′ end and simultaneously joins the two exons. The intron lariat is released and subsequently debranched and degraded. The spliced mRNA is then exported from the nucleus for translation.

3 Splice site sequences and consensus patterns

Splice site sequences are not identical across all introns; they follow degenerate consensus patterns that allow varied binding affinity and regulation.

3.1 Consensus sequences in vertebrates, plants, and yeast

In vertebrates, the 5′ splice site consensus is AGGURAGU, and the 3′ splice site is (Y)nNYAGG. In plants, the 5′ splice site consensus is similar but often less strict (e.g., AGGUAAGU). In the yeast *Saccharomyces cerevisiae*, the 5′ splice site is almost invariant (GUAUGU) and the 3′ splice site is YAG, with a highly conserved branch point sequence (UACUAAC). These differences reflect the relative simplicity of yeast splicing versus the complex regulation in metazoans.

3.2 Degeneracy and splice site strength

Splice site sequences vary in how well they match the consensus. A close match yields a “strong” splice site, typically used efficiently in constitutive splicing. Poor matches yield “weak” splice sites, which are often targets for regulation via alternative splicing. The strength of a splice site influences the kinetics of spliceosome assembly and can be modulated by auxiliary splicing factors.

3.3 Alternative splicing and splice site selection

Alternative splicing is a process by which a single pre‑mRNA can generate multiple mRNA isoforms through differential use of splice sites. This greatly expands the proteome and allows tissue‑specific or developmental regulation.

3.3.1 Exon skipping

In exon skipping, an exon is excluded from the mature mRNA. Both the upstream and downstream introns are spliced out, joining the flanking exons. This is the most common form of alternative splicing in mammals. It is often controlled by regulatory sequences (exonic splicing enhancers or silencers) and their binding proteins.

3.3.2 Alternative 5′ or 3′ splice sites

Alternative 5′ splice sites use a different 5′ splice site within the same exon or intron, leading to exon extension or truncation. Similarly, alternative 3′ splice sites change the downstream boundary. This can alter the protein sequence or introduce premature stop codons.

3.3.3 Intron retention

Intron retention occurs when an intron is not removed and remains in the mature mRNA. This often leads to nonsense‑mediated decay or production of a truncated protein. Intron retention is a common regulatory mechanism in plants and some metazoans.

4 Splice site mutations and disease

Mutations that alter splice site sequences can disrupt normal splicing, leading to aberrant mRNA isoforms and often causing genetic disorders.

4.1 Mechanism of splicing disruption

Splice site mutations can weaken or strengthen splice sites, create cryptic splice sites (previously unused sites that become active), or abolish splicing entirely. Common types include point mutations in the invariant GU or AG dinucleotides, which typically abolish splicing at that site, and mutations in the branch point or polypyrimidine tract, which reduce splicing efficiency. Such mutations may lead to exon skipping, intron retention, or use of cryptic splice sites, often resulting in frameshifts or inclusion of premature termination codons.

4.2 Examples of splice site mutations in human genetic disorders

Many human diseases are caused by splice site mutations. For example, mutations in the 5′ splice site of the β‑globin gene cause β‑thalassemia. In spinal muscular atrophy, a silent mutation in the SMN2 gene alters an exonic splicing enhancer, leading to exon 7 skipping. Mutations in the fibrillin‑1 gene (FBN1) causing Marfan syndrome frequently affect splice sites. Cystic fibrosis can result from splice site mutations in the CFTR gene. These examples illustrate the broad impact of splicing defects.

4.3 Splice‑switching therapeutics (antisense oligonucleotides)

Antisense oligonucleotides (ASOs) can be designed to bind specific pre‑mRNA regions and modulate splicing. For splice site mutations, ASOs can block cryptic splice sites or enhance inclusion of skipped exons. For example, nusinersen (Spinraza) is an ASO approved for spinal muscular atrophy; it promotes inclusion of SMN2 exon 7 by blocking an intronic silencer. Other splice‑switching ASOs are in clinical trials for Duchenne muscular dystrophy and other disorders.

5 Bioinformatic prediction and analysis of splice sites

Computational tools are essential for identifying splice sites in genomic sequences and predicting the effects of mutations.

5.1 Position weight matrices and scoring

Position weight matrices (PWMs) are the classical method for splice site prediction. A PWM is derived from aligned sequences of known splice sites, assigning a score for each nucleotide at each position. The score of a candidate site is the sum of the log‑odds ratios; higher scores indicate better matches to the consensus. Tools like MaxEntScan use maximum entropy models to improve accuracy over simple PWMs.

5.2 Machine learning approaches (deep learning, SVM)

More recent methods employ machine learning. Support vector machines (SVMs) use features such as sequence composition and conservation to classify sites. Deep learning models (e.g., SpliceAI, DeepSplice) use convolutional or recurrent neural networks trained on large datasets of splice sites. These methods achieve high accuracy and can predict the effect of single‑nucleotide variants on splicing.

5.3 Database resources (SpliceDB, ENSEMBL splice site annotations)

Several databases compile known splice sites. SpliceDB contains vertebrate splice site sequences. ENSEMBL provides genome annotations that include exon‑intron boundaries and alternative splicing patterns. The UCSC Genome Browser also displays splice site annotations and conservation tracks. These resources are used for both research and clinical variant interpretation.

6 Experimental methods for splice site identification

Experimental approaches are used to validate predicted splice sites and characterize splicing patterns.

6.1 RT‑PCR and sequencing of spliced products

Reverse transcription polymerase chain reaction (RT‑PCR) amplifies cDNA from RNA, allowing detection of splice isoforms. Gel electrophoresis or capillary electrophoresis reveals the sizes of products, and sequencing confirms exact exon junctions. This method is widely used to assess the effects of mutations on splicing.

6.2 Minigene splicing reporter assays

Minigene assays involve cloning a genomic region of interest (including exons and introns) into a reporter vector, transfecting it into cells, and analyzing the spliced RNA. The minigene provides a simplified system to test how mutations or trans‑acting factors affect splice site usage. Often the reporter includes a fluorescent or luminescent protein to quantify splicing efficiency.

6.3 High‑throughput methods (RNA‑seq, CLIP‑seq)

RNA‑sequencing (RNA‑seq) provides genome‑wide information on splicing patterns. Short reads are aligned to the genome to detect splice junctions. Advanced analysis can quantify exon inclusion levels and detect novel splice sites. Crosslinking immunoprecipitation sequencing (CLIP‑seq) identifies RNA‑binding protein targets and can map regulatory elements that influence splice site selection. High‑throughput methods have greatly accelerated the discovery of splicing regulation and disease‑associated variants.