Here are some widely used chemical compound databases, organized by purpose. I’ve included what they’re best for and whether they’re free or typically paid.
General chemical structures and properties (free or widely accessible)
- PubChem (NIH/NCBI) — Huge free database of chemical structures, properties, synonyms, and links to bioassay data.
- ChemSpider (Royal Society of Chemistry) — Aggregated public sources; structure, properties, and multiple ways to search (including substructure).
- ChEBI (EMBL-EBI) — Curated ontology of chemical entities of biological interest; great for standardized naming and classifications.
- ZINC — Free subset for purchasable compounds; nice for virtual screening with ready-made 3D conformations.
- Lipid Maps Structure Database (LMSD) — For lipid structures and related data.
Drug-focused and pharmacology data
- DrugBank — Detailed drug data: chemical structures, pharmacology, mechanisms, interactions; some content free, more available with registration.
- KEGG DRUG — Curated drugs with structures, classifications, and pathway context.
- ChEMBL — Curated bioactivity database for drug-like molecules; many target-assay relationships.
- BindingDB — Binding affinities (Kd, Ki, IC50) for compounds against targets; useful for SAR and SAR-to-target links.
Natural products and metabolites
- NPASS (Natural Product Activity and Species Source Database) — Natural products linked to activities and sources.
- COCONUT (Natural Products Atlas) — Open, large collection of natural product structures.
- HMDB (Human Metabolome Database) — Human metabolites with structure, properties, and spectra.
- METLIN (metabolite and MS data) — Metabolite spectra and identities (frequently subscription-based, some free content).
Structure and drug-like filtering for discovery
- ZINC (also listed above) — Special emphasis on purchasable compounds for docking and screening; good 3D geometry and filtering.
- eMolecules / Enamine (catalog catalogs) — Commercial catalogs; useful for procurement and property filtering (usually require user accounts or subscriptions).
Structural biology and 3D structures
- Protein Data Bank (PDB) — 3D structures of proteins, nucleic acids, complexes (not strictly “small molecules” but critical for structure-based design).
- PDBe / PDBe-KB — Integrated structural data with annotations; links to ligands and binding sites.
- AlphaFold Protein Structure Database — Predicted protein structures with confidence scores.
Literature and broader context (often paywalled, but essential)
- SciFinder, Reaxys — Comprehensive but typically require institutional/licensed access; excellent for literature-linked chemistry and reactions.
- TOXNET/NLM resources — Toxicology data and chemical safety information (some content integrated into other NLM databases).
Tips to choose
- For quick chemical lookup and basic properties: PubChem or ChemSpider.
- For drug discovery and pharmacology: DrugBank, ChEMBL, BindingDB.
- For natural products or metabolites: NPASS, COCONUT, HMDB.
- For structural biology context: PDB / PDBe, plus AlphaFold for proteins without experimental structures.
- For procurement or virtual screening: ZINC (and Enamine/eMolecules catalogs).
- For strict ontology and standard naming: ChEBI.
If you tell me what you’re researching (drug discovery, metabolomics, natural products catalog, structure-based design, etc.), I can tailor a short list with direct links, and suggest search queries and API options.