Drug discovery is the process by which new therapeutic molecules are identified, designed, and developed to treat or prevent diseases. It integrates multiple scientific disciplines—including medicinal chemistry, pharmacology, molecular biology, and computational modeling—to transform a biological target or disease hypothesis into a safe and effective drug candidate. The journey from initial concept to approved medication typically takes 10–15 years and involves iterative cycles of hypothesis testing, lead optimization, and preclinical validation before entering human clinical trials.
1 Historical background
1.1 Traditional approaches (serendipity and natural products)
Before the 20th century, most medicines were discovered by chance or derived from folklore remedies. Plant extracts, minerals, and animal products were used empirically. For example, digitalis from foxglove was used for heart conditions, and quinine from cinchona bark treated malaria. Early pharmacology relied on whole-organism effects, and active principles were rarely isolated.
1.2 Emergence of rational drug design
In the late 19th and early 20th centuries, advances in organic chemistry and physiology enabled a more systematic approach. Paul Ehrlich’s concept of the “magic bullet” led to the first synthetic drugs, such as arsphenamine for syphilis. Structure–activity relationships began to be explored, and by the mid-20th century, drug discovery incorporated receptor theory and enzyme inhibition.
1.3 Modern era: high-throughput screening and genomics
The 1980s and 1990s saw automation and miniaturization of assays, allowing millions of compounds to be tested rapidly against a target. The Human Genome Project (completed 2003) identified thousands of potential new drug targets. Combinatorial chemistry and later DNA-encoded libraries expanded chemical space. Today, genomics, proteomics, and systems biology inform target selection.
2 Core stages of drug discovery
2.1 Target identification and validation
2.1.1 Disease association and genetic evidence
A drug target is a biomolecule (often a protein) whose modulation is expected to alter disease course. Genetic studies—such as genome-wide association studies (GWAS), familial linkage, or CRISPR-based screens—provide evidence linking a target to a disease. For example, mutations in the *PCSK9* gene were linked to cholesterol levels, leading to a new class of lipid-lowering drugs.
2.1.2 Target druggability assessment
Not all biologically relevant targets are amenable to small-molecule modulation. Druggability is assessed by examining the target’s structure, binding site properties, and known ligandability. Structural biology (e.g., X-ray crystallography) and computational analyses help predict whether a target can accommodate a drug-like molecule. Targets with shallow or highly polar binding sites are often considered challenging.
2.2 Hit discovery
2.2.1 High-throughput screening (HTS)
HTS uses automated liquid handlers and microplates to test hundreds of thousands of compounds in biochemical or cellular assays. A “hit” is a compound that shows reproducible activity above a predefined threshold. Modern HTS libraries can contain millions of diverse molecules, and advanced detection methods (fluorescence, luminescence, mass spectrometry) enable rapid data collection.
2.2.2 Fragment-based screening
Fragment-based drug discovery (FBDD) uses low-molecular-weight compounds (typically <300 Da) that bind weakly to the target. Hits are detected by biophysical methods such as NMR, surface plasmon resonance (SPR), or X-ray crystallography. Fragments are then merged or grown into larger leads with improved potency, often guided by structural information.
2.2.3 Virtual screening and computational methods
2.2.3.1 Molecular docking
Docking algorithms predict the preferred orientation and binding affinity of a small molecule within a target’s binding site. Large virtual libraries can be screened in silico, ranking compounds by predicted binding energy. Docking is widely used to identify novel scaffolds, but accuracy depends on the quality of the target structure and scoring functions.
2.2.3.2 Pharmacophore modeling
A pharmacophore is a spatial arrangement of chemical features (hydrogen bond donors/acceptors, hydrophobic regions, charged groups) essential for biological activity. Computational models are built from known active compounds or from target–ligand complexes. Virtual screening then retrieves molecules that match the pharmacophore, enriching for potential hits.
2.3 Hit-to-lead and lead optimization
2.3.1 Structure-activity relationship (SAR) studies
SAR systematically explores how changes in chemical structure affect biological activity. Iterative cycles of synthesis and testing generate data that guide optimization. Quantitative SAR (QSAR) uses statistical models to correlate molecular descriptors with potency, enabling predictive design.
2.3.2 Medicinal chemistry strategies
Medicinal chemists employ various tactics to improve potency, selectivity, and drug-like properties. These include scaffold hopping, isosteric replacement, introduction of conformational constraints, and modification of functional groups to reduce metabolism or enhance solubility. The goal is to produce a lead compound with a balanced profile.
2.3.3 ADME-Tox profiling
2.3.3.1 Absorption and bioavailability
Absorption, distribution, metabolism, excretion, and toxicity (ADME-Tox) are assessed early to avoid late-stage failures. Permeability (e.g., Caco-2 cell assays) and solubility (kinetic and thermodynamic) predict oral bioavailability. Efflux transporters (e.g., P-glycoprotein) can limit brain penetration or intestinal absorption.
2.3.3.2 Metabolism and excretion
Metabolic stability is measured using liver microsomes or hepatocytes. Cytochrome P450 enzymes often mediate oxidative metabolism. Identification of major metabolites and their activities is important. Excretion routes (renal, biliary) are studied in preclinical species to predict human clearance.
2.3.3.3 Toxicity screening (in vitro and in vivo)
In vitro assays detect cytotoxicity, genotoxicity (Ames test), hERG channel blockade (cardiac risk), and phospholipidosis. In vivo rodent studies assess acute and repeated-dose toxicity. These data help select candidates with a favorable safety margin before clinical trials.
2.4 Preclinical development
2.4.1 Pharmacokinetics and pharmacodynamics
Pharmacokinetics (PK) describes the time course of drug concentration in the body (absorption, distribution, metabolism, excretion). Pharmacodynamics (PD) relates concentration to effect. Integrated PK/PD modeling determines dosing regimens and predicts human exposure. Studies are performed in rodents, dogs, or non-human primates.
2.4.2 Safety pharmacology and toxicology
Safety pharmacology evaluates effects on vital organ systems (cardiovascular, respiratory, central nervous system). Toxicology studies include single-dose, repeated-dose, and reproductive toxicity tests. Genotoxicity and carcinogenicity assessments are also required. Good Laboratory Practice (GLP) governs these studies.
2.4.3 Formulation and route of administration
Formulation development aims to deliver the drug to the target site with acceptable stability, bioavailability, and patient compliance. Oral formulations (tablets, capsules) are common, but parenteral, inhaled, or transdermal routes may be used. Excipients, pH, particle size, and release characteristics are optimized.
3 Key enabling technologies
3.1 Computational chemistry and AI
3.1.1 Machine learning for property prediction
Machine learning models trained on large datasets predict ADME-Tox properties, solubility, permeability, and off-target effects. Deep learning (e.g., graph neural networks) can process molecular graphs directly. These predictions prioritize compounds for synthesis and reduce experimental burden.
3.1.2 De novo drug design
Generative models (e.g., variational autoencoders, generative adversarial networks, reinforcement learning) propose novel molecular structures that satisfy multiple constraints (target affinity, synthesizability, drug-likeness). De novo design explores chemical space beyond existing libraries and can produce patentable scaffolds.
3.2 Structural biology
3.2.1 X-ray crystallography and cryo-EM
X-ray crystallography provides atomic-resolution structures of target–ligand complexes, enabling structure-based drug design. Cryo-electron microscopy (cryo-EM) has become a powerful tool for large complexes and membrane proteins that are difficult to crystallize, such as GPCRs and ion channels.
3.2.2 NMR spectroscopy
Nuclear magnetic resonance (NMR) is used for fragment screening, studying protein–ligand interactions in solution, and determining structures of small proteins or domains. Methods like saturation transfer difference (STD) and waterLOGSY detect weak binding events. NMR also provides dynamic information on conformational changes.
3.3 Biophysical and biochemical assays
3.3.1 Binding affinity assays
Biophysical techniques quantify binding affinity (Kd, Ki, IC50) without requiring functional activity. Common methods include SPR, isothermal titration calorimetry (ITC), and fluorescence polarization. These assays confirm direct target engagement and help rank compounds.
3.3.2 Functional cell-based assays
Cell-based assays measure functional responses such as second messenger levels, reporter gene expression, or cell proliferation. They assess target engagement in a physiological context, detect off-target effects, and provide early evidence of efficacy. High-content imaging and flow cytometry add spatial information.
4 Special areas
4.1 Natural product-based discovery
Natural products continue to inspire drug leads, especially for antibiotics and anticancer agents. Modern approaches include bioactivity-guided fractionation, genome mining for biosynthetic gene clusters, and semi-synthetic modification of known scaffolds. Challenges include supply, complex structures, and intellectual property.
4.2 Biologics and peptide drugs
4.2.1 Monoclonal antibodies
Monoclonal antibodies (mAbs) are produced by hybridoma or recombinant technologies. They target extracellular proteins with high specificity. Advances in humanization, phage display, and transgenic mice have reduced immunogenicity. IgG formats are predominant, but fragments (Fab, scFv) and antibody–drug conjugates (ADCs) are also used.
4.2.2 Macrocyclic peptides
Macrocyclic peptides combine high potency and target selectivity with oral bioavailability in some cases. They are discovered by phage display, mRNA display, or chemical synthesis. Cyclosporine A (immunosuppressant) is a classic example. Strategies to improve permeability and metabolic stability are active research areas.
4.3 Repurposing and polypharmacology
Drug repurposing identifies new indications for existing drugs, reducing development costs and time. Systematic approaches include computational similarity searches, phenotypic screening, and retrospective clinical data analysis. Polypharmacology aims to design compounds that modulate multiple targets, mimicking combination therapy.
4.4 RNA-targeted and antisense therapies
RNA-based drugs include antisense oligonucleotides (ASOs) that bind to mRNA and modulate splicing or degrade transcripts, small interfering RNAs (siRNAs) that trigger RNA interference, and small molecules that target structured RNA (e.g., riboswitches). GalNAc conjugation improves liver delivery. Several ASOs and siRNAs are approved for genetic disorders.
5 Regulatory and translational aspects
5.1 IND-enabling studies
Before a candidate can be tested in humans, an Investigational New Drug (IND) application must be filed with regulatory agencies (e.g., FDA, EMA). It includes data on chemistry, manufacturing, controls (CMC), pharmacology, toxicology, and proposed clinical protocols. Regulatory review typically takes 30 days.
5.2 Clinical trial phases (I–III)
Phase I trials assess safety, tolerability, and pharmacokinetics in a small number (20–100) of healthy volunteers or patients. Phase II evaluates efficacy and dose–response in several hundred patients with the target disease. Phase III is a large-scale (hundreds to thousands) randomized controlled trial to confirm efficacy and monitor adverse events. Success in Phase III leads to a New Drug Application (NDA) or Marketing Authorization Application (MAA).
5.3 Post-approval and pharmacovigilance
After approval, Phase IV studies (post-marketing surveillance) monitor long-term safety and real-world effectiveness. Pharmacovigilance systems collect adverse event reports, and regulators can require label changes or withdrawals if new risks emerge.
6 Ethical and economic considerations
6.1 Animal testing alternatives
Regulatory requirements for animal testing are increasingly questioned. Alternatives include in vitro models (organ-on-a-chip, 3D cell cultures), computational modeling, and human microdosing studies. The 3Rs (Replacement, Reduction, Refinement) guide ethical animal use, but complete replacement remains challenging due to whole-body complexities.
6.2 Cost and risk of attrition
Developing a new drug costs an estimated $1–2 billion, mostly due to high attrition rates. Only about 10% of candidates entering Phase I eventually receive approval. Failures often result from lack of efficacy, toxicology issues, or poor pharmacokinetics. Strategies like biomarker-driven patient selection and adaptive trial designs aim to reduce attrition.
6.3 Access and affordability
High drug prices raise equity concerns, especially for rare diseases or global health priorities. Pricing debates involve research costs, patent monopolies, and insurance systems. Efforts to improve access include tiered pricing, compulsory licensing, and public–private partnerships (e.g., Medicines Patent Pool).