rdkit ยท diff

v1.1 to v1.2

21 added, 757 removed. Audit A to A.

---
name: rdkit
description: Cheminformatics toolkit for fine-grained molecular control. SMILES/SDF parsing, descriptors (MW, LogP, TPSA), fingerprints, substructure search, 2D/3D generation, similarity, reactions. For standard workflows with simpler interface, use datamol (wrapper around RDKit). Use rdkit for advanced control, custom sanitization, specialized algorithms.
license: BSD-3-Clause license
allowed-tools: Read Write Edit Bash
compatibility: Examples target RDKit 2026.03.x. Use conda-forge for the broadest binary support or PyPI package `rdkit` for supported platform wheels; `rdkit-pypi` is the legacy PyPI name.
metadata:
- version: "1.1"
+ version: "1.2"
skill-author: K-Dense Inc.
---
# RDKit Cheminformatics Toolkit
## Overview
RDKit is a comprehensive cheminformatics library providing Python APIs for molecular analysis and manipulation. This skill provides guidance for reading/writing molecular structures, calculating descriptors, fingerprinting, substructure searching, chemical reactions, 2D/3D coordinate generation, and molecular visualization. Use this skill for drug discovery, computational chemistry, and cheminformatics research tasks.
**Current baseline (checked 2026-06-07):** RDKit **2026.03.3** is the latest GitHub/PyPI release (`rdkit` 2026.3.3 on PyPI). Official installation docs continue to recommend conda-forge for most users, while cross-platform PyPI wheels are published under the `rdkit` package name. `rdkit-pypi` is the old PyPI package name and should only appear when maintaining legacy environments.
## Installation and Setup
Use `uv` when installing into an existing Python environment:
```bash
uv pip install rdkit
```
For reproducible chemistry environments, especially when mixing compiled scientific packages, conda-forge remains the upstream recommendation:
```bash
conda create -c conda-forge -n my-rdkit-env rdkit
conda activate my-rdkit-env
```
Avoid installing both conda `rdkit` and PyPI `rdkit`/`rdkit-pypi` into the same environment unless you are deliberately debugging packaging behavior. Mixed installs can make it unclear which binary extension is being imported.
## Core Capabilities
- ### 1. Molecular I/O and Creation
-
- **Reading Molecules:**
-
- Read molecular structures from various formats:
-
- ```python
- from rdkit import Chem
-
- # From SMILES strings
- mol = Chem.MolFromSmiles('Cc1ccccc1') # Returns Mol object or None
-
- # From MOL files
- mol = Chem.MolFromMolFile('path/to/file.mol')
-
- # From MOL blocks (string data)
- mol = Chem.MolFromMolBlock(mol_block_string)
-
- # From InChI
- mol = Chem.MolFromInchi('InChI=1S/C6H6/c1-2-4-6-5-3-1/h1-6H')
- ```
-
- **Writing Molecules:**
-
- Convert molecules to text representations:
-
- ```python
- # To canonical SMILES
- smiles = Chem.MolToSmiles(mol)
-
- # To MOL block
- mol_block = Chem.MolToMolBlock(mol)
-
- # To InChI
- inchi = Chem.MolToInchi(mol)
- ```
-
- **Batch Processing:**
-
- For processing multiple molecules, use Supplier/Writer objects:
-
- ```python
- # Read SDF files
- suppl = Chem.SDMolSupplier('molecules.sdf')
- for mol in suppl:
- if mol is not None: # Check for parsing errors
- # Process molecule
- pass
-
- # Read SMILES files
- suppl = Chem.SmilesMolSupplier('molecules.smi', titleLine=False)
-
- # For large files or compressed data
- import gzip
-
- with gzip.open('molecules.sdf.gz') as f:
- suppl = Chem.ForwardSDMolSupplier(f)
- for mol in suppl:
- # Process molecule
- pass
-
- # Multithreaded processing for large datasets
- suppl = Chem.MultithreadedSDMolSupplier('molecules.sdf')
-
- # Write molecules to SDF
- writer = Chem.SDWriter('output.sdf')
- for mol in molecules:
- writer.write(mol)
- writer.close()
- ```
-
- **Important Notes:**
- - All `MolFrom*` functions return `None` on failure with error messages
- - Always check for `None` before processing molecules
- - Molecules are automatically sanitized on import (validates valence, perceives aromaticity)
-
- ### 2. Molecular Sanitization and Validation
-
- RDKit automatically sanitizes molecules during parsing, executing 13 steps including valence checking, aromaticity perception, and chirality assignment.
-
- **Sanitization Control:**
-
- ```python
- # Disable automatic sanitization
- mol = Chem.MolFromSmiles('C1=CC=CC=C1', sanitize=False)
-
- # Manual sanitization
- Chem.SanitizeMol(mol)
-
- # Detect problems before sanitization
- problems = Chem.DetectChemistryProblems(mol)
- for problem in problems:
- print(problem.GetType(), problem.Message())
-
- # Partial sanitization (skip specific steps)
- Chem.SanitizeMol(mol, sanitizeOps=Chem.SANITIZE_ALL ^ Chem.SANITIZE_PROPERTIES)
- ```
-
- **Common Sanitization Issues:**
- - Atoms with explicit valence exceeding maximum allowed will raise exceptions
- - Invalid aromatic rings will cause kekulization errors
- - Radical electrons may not be properly assigned without explicit specification
-
- ### 3. Molecular Analysis and Properties
-
- **Accessing Molecular Structure:**
-
- ```python
- # Iterate atoms and bonds
- for atom in mol.GetAtoms():
- print(atom.GetSymbol(), atom.GetIdx(), atom.GetDegree())
-
- for bond in mol.GetBonds():
- print(bond.GetBeginAtomIdx(), bond.GetEndAtomIdx(), bond.GetBondType())
-
- # Ring information
- ring_info = mol.GetRingInfo()
- ring_info.NumRings()
- ring_info.AtomRings() # Returns tuples of atom indices
-
- # Check if atom is in ring
- atom = mol.GetAtomWithIdx(0)
- atom.IsInRing()
- atom.IsInRingSize(6) # Check for 6-membered rings
-
- # Find smallest set of smallest rings (SSSR)
- from rdkit.Chem import GetSymmSSSR
- rings = GetSymmSSSR(mol)
- ```
-
- **Stereochemistry:**
-
- ```python
- # Find chiral centers
- from rdkit.Chem import FindMolChiralCenters
- chiral_centers = FindMolChiralCenters(mol, includeUnassigned=True)
- # Returns list of (atom_idx, chirality) tuples
-
- # Assign stereochemistry from 3D coordinates
- from rdkit.Chem import AssignStereochemistryFrom3D
- AssignStereochemistryFrom3D(mol)
-
- # Check bond stereochemistry
- bond = mol.GetBondWithIdx(0)
- stereo = bond.GetStereo() # STEREONONE, STEREOZ, STEREOE, etc.
- ```
-
- **Fragment Analysis:**
-
- ```python
- # Get disconnected fragments
- frags = Chem.GetMolFrags(mol, asMols=True)
-
- # Fragment on specific bonds
- from rdkit.Chem import FragmentOnBonds
- frag_mol = FragmentOnBonds(mol, [bond_idx1, bond_idx2])
-
- # Count ring systems
- from rdkit.Chem.Scaffolds import MurckoScaffold
- scaffold = MurckoScaffold.GetScaffoldForMol(mol)
- ```
-
- ### 4. Molecular Descriptors and Properties
-
- **Basic Descriptors:**
-
- ```python
- from rdkit.Chem import Descriptors
-
- # Molecular weight
- mw = Descriptors.MolWt(mol)
- exact_mw = Descriptors.ExactMolWt(mol)
-
- # LogP (lipophilicity)
- logp = Descriptors.MolLogP(mol)
-
- # Topological polar surface area
- tpsa = Descriptors.TPSA(mol)
-
- # Number of hydrogen bond donors/acceptors
- hbd = Descriptors.NumHDonors(mol)
- hba = Descriptors.NumHAcceptors(mol)
-
- # Number of rotatable bonds
- rot_bonds = Descriptors.NumRotatableBonds(mol)
-
- # Number of aromatic rings
- aromatic_rings = Descriptors.NumAromaticRings(mol)
- ```
-
- **Batch Descriptor Calculation:**
-
- ```python
- # Calculate all descriptors at once
- all_descriptors = Descriptors.CalcMolDescriptors(mol)
- # Returns dictionary: {'MolWt': 180.16, 'MolLogP': 1.23, ...}
-
- # Get list of available descriptor names
- descriptor_names = [desc[0] for desc in Descriptors._descList]
- ```
-
- **Lipinski's Rule of Five:**
-
- ```python
- # Check drug-likeness
- mw = Descriptors.MolWt(mol) <= 500
- logp = Descriptors.MolLogP(mol) <= 5
- hbd = Descriptors.NumHDonors(mol) <= 5
- hba = Descriptors.NumHAcceptors(mol) <= 10
-
- is_drug_like = mw and logp and hbd and hba
- ```
-
- ### 5. Fingerprints and Molecular Similarity
-
- **Fingerprint Types:**
-
- ```python
- from rdkit.Chem import rdFingerprintGenerator
- from rdkit.Chem import MACCSkeys
-
- # RDKit topological fingerprint
- rdk_gen = rdFingerprintGenerator.GetRDKitFPGenerator(minPath=1, maxPath=7, fpSize=2048)
- fp = rdk_gen.GetFingerprint(mol)
-
- # Morgan fingerprints (circular fingerprints, similar to ECFP)
- # Modern API using rdFingerprintGenerator
- morgan_gen = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
- fp = morgan_gen.GetFingerprint(mol)
- # Count-based fingerprint
- fp_count = morgan_gen.GetCountFingerprint(mol)
-
- # MACCS keys (166-bit structural key)
- fp = MACCSkeys.GenMACCSKeys(mol)
-
- # Atom pair fingerprints
- ap_gen = rdFingerprintGenerator.GetAtomPairGenerator()
- fp = ap_gen.GetFingerprint(mol)
-
- # Topological torsion fingerprints
- tt_gen = rdFingerprintGenerator.GetTopologicalTorsionGenerator()
- fp = tt_gen.GetFingerprint(mol)
-
- # Avalon fingerprints (if available)
- from rdkit.Avalon import pyAvalonTools
- fp = pyAvalonTools.GetAvalonFP(mol)
- ```
-
- **Similarity Calculation:**
-
- ```python
- from rdkit import DataStructs
- from rdkit.Chem import rdFingerprintGenerator
-
- # Generate fingerprints using generator
- mfpgen = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
- fp1 = mfpgen.GetFingerprint(mol1)
- fp2 = mfpgen.GetFingerprint(mol2)
-
- # Calculate Tanimoto similarity
- similarity = DataStructs.TanimotoSimilarity(fp1, fp2)
-
- # Calculate similarity for multiple molecules
- fps = [mfpgen.GetFingerprint(m) for m in [mol2, mol3, mol4]]
- similarities = DataStructs.BulkTanimotoSimilarity(fp1, fps)
-
- # Other similarity metrics
- dice = DataStructs.DiceSimilarity(fp1, fp2)
- cosine = DataStructs.CosineSimilarity(fp1, fp2)
- ```
-
- **Clustering and Diversity:**
-
- ```python
- # Butina clustering based on fingerprint similarity
- from rdkit.ML.Cluster import Butina
-
- # Calculate distance matrix
- dists = []
- mfpgen = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
- fps = [mfpgen.GetFingerprint(mol) for mol in mols]
- for i in range(len(fps)):
- sims = DataStructs.BulkTanimotoSimilarity(fps[i], fps[:i])
- dists.extend([1-sim for sim in sims])
-
- # Cluster with distance cutoff
- clusters = Butina.ClusterData(dists, len(fps), distThresh=0.3, isDistData=True)
- ```
-
- ### 6. Substructure Searching and SMARTS
-
- **Basic Substructure Matching:**
-
- ```python
- # Define query using SMARTS
- query = Chem.MolFromSmarts('[#6]1:[#6]:[#6]:[#6]:[#6]:[#6]:1') # Benzene ring
-
- # Check if molecule contains substructure
- has_match = mol.HasSubstructMatch(query)
-
- # Get all matches (returns tuple of tuples with atom indices)
- matches = mol.GetSubstructMatches(query)
-
- # Get only first match
- match = mol.GetSubstructMatch(query)
- ```
-
- **Common SMARTS Patterns:**
-
- ```python
- # Primary alcohols
- primary_alcohol = Chem.MolFromSmarts('[CH2][OH1]')
-
- # Carboxylic acids
- carboxylic_acid = Chem.MolFromSmarts('C(=O)[OH]')
-
- # Amides
- amide = Chem.MolFromSmarts('C(=O)N')
-
- # Aromatic heterocycles
- aromatic_n = Chem.MolFromSmarts('[nR]') # Aromatic nitrogen in ring
-
- # Macrocycles (rings > 12 atoms)
- macrocycle = Chem.MolFromSmarts('[r{12-}]')
- ```
-
- **Matching Rules:**
- - Unspecified properties in query match any value in target
- - Hydrogens are ignored unless explicitly specified
- - Charged query atom won't match uncharged target atom
- - Aromatic query atom won't match aliphatic target atom (unless query is generic)
-
- ### 7. Chemical Reactions
-
- **Reaction SMARTS:**
-
- ```python
- from rdkit.Chem import AllChem
-
- # Define reaction using SMARTS: reactants >> products
- rxn = AllChem.ReactionFromSmarts('[C:1]=[O:2]>>[C:1][O:2]') # Ketone reduction
-
- # Apply reaction to molecules
- reactants = (mol1,)
- products = rxn.RunReactants(reactants)
-
- # Products is tuple of tuples (one tuple per product set)
- for product_set in products:
- for product in product_set:
- # Sanitize product
- Chem.SanitizeMol(product)
- ```
-
- **Reaction Features:**
- - Atom mapping preserves specific atoms between reactants and products
- - Dummy atoms in products are replaced by corresponding reactant atoms
- - "Any" bonds inherit bond order from reactants
- - Chirality preserved unless explicitly changed
-
- **Reaction Similarity:**
-
- ```python
- # Generate reaction fingerprints
- fp = AllChem.CreateDifferenceFingerprintForReaction(rxn)
-
- # Compare reactions
- similarity = DataStructs.TanimotoSimilarity(fp1, fp2)
- ```
-
- ### 8. 2D and 3D Coordinate Generation
-
- **2D Coordinate Generation:**
-
- ```python
- from rdkit.Chem import AllChem
-
- # Generate 2D coordinates for depiction
- AllChem.Compute2DCoords(mol)
-
- # Align molecule to template structure
- template = Chem.MolFromSmiles('c1ccccc1')
- AllChem.Compute2DCoords(template)
- AllChem.GenerateDepictionMatching2DStructure(mol, template)
- ```
-
- **3D Coordinate Generation and Conformers:**
-
- ```python
- # Generate single 3D conformer using ETKDG
- AllChem.EmbedMolecule(mol, randomSeed=42)
-
- # Generate multiple conformers
- conf_ids = AllChem.EmbedMultipleConfs(mol, numConfs=10, randomSeed=42)
-
- # Optimize geometry with force field
- AllChem.UFFOptimizeMolecule(mol) # UFF force field
- AllChem.MMFFOptimizeMolecule(mol) # MMFF94 force field
-
- # Optimize all conformers
- for conf_id in conf_ids:
- AllChem.MMFFOptimizeMolecule(mol, confId=conf_id)
-
- # Calculate RMSD between conformers
- from rdkit.Chem import AllChem
- rms = AllChem.GetConformerRMS(mol, conf_id1, conf_id2)
-
- # Align molecules
- AllChem.AlignMol(probe_mol, ref_mol)
- ```
-
- **Constrained Embedding:**
-
- ```python
- # Embed with part of molecule constrained to specific coordinates
- AllChem.ConstrainedEmbed(mol, core_mol)
- ```
-
- ### 9. Molecular Visualization
-
- **Basic Drawing:**
-
- ```python
- from rdkit.Chem import Draw
-
- # Draw single molecule to PIL image
- img = Draw.MolToImage(mol, size=(300, 300))
- img.save('molecule.png')
-
- # Draw to file directly
- Draw.MolToFile(mol, 'molecule.png')
-
- # Draw multiple molecules in grid
- mols = [mol1, mol2, mol3, mol4]
- img = Draw.MolsToGridImage(mols, molsPerRow=2, subImgSize=(200, 200))
- ```
-
- **Highlighting Substructures:**
-
- ```python
- # Highlight substructure match
- query = Chem.MolFromSmarts('c1ccccc1')
- match = mol.GetSubstructMatch(query)
-
- img = Draw.MolToImage(mol, highlightAtoms=match)
-
- # Custom highlight colors
- highlight_colors = {atom_idx: (1, 0, 0) for atom_idx in match} # Red
- img = Draw.MolToImage(mol, highlightAtoms=match,
- highlightAtomColors=highlight_colors)
- ```
-
- **Customizing Visualization:**
-
- ```python
- from rdkit.Chem.Draw import rdMolDraw2D
-
- # Create drawer with custom options
- drawer = rdMolDraw2D.MolDraw2DCairo(300, 300)
- opts = drawer.drawOptions()
-
- # Customize options
- opts.addAtomIndices = True
- opts.addStereoAnnotation = True
- opts.bondLineWidth = 2
-
- # Draw molecule
- drawer.DrawMolecule(mol)
- drawer.FinishDrawing()
-
- # Save to file
- with open('molecule.png', 'wb') as f:
- f.write(drawer.GetDrawingText())
- ```
-
- **Jupyter Notebook Integration:**
-
- ```python
- # Enable inline display in Jupyter
- from rdkit.Chem.Draw import IPythonConsole
-
- # Customize default display
- IPythonConsole.ipython_useSVG = True # Use SVG instead of PNG
- IPythonConsole.molSize = (300, 300) # Default size
-
- # Molecules now display automatically
- mol # Shows molecule image
- ```
-
- **Visualizing Fingerprint Bits:**
-
- ```python
- # Show what molecular features a fingerprint bit represents
- from rdkit.Chem import Draw
- from rdkit.Chem import rdFingerprintGenerator
-
- # For Morgan fingerprints
- additional_output = rdFingerprintGenerator.AdditionalOutput()
- additional_output.AllocateBitInfoMap()
- morgan_gen = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
- fp = morgan_gen.GetFingerprint(mol, additionalOutput=additional_output)
- bit_info = additional_output.GetBitInfoMap()
-
- # Draw environment for specific bit
- img = Draw.DrawMorganBit(mol, bit_id, bit_info)
- ```
-
- ### 10. Molecular Modification
-
- **Adding/Removing Hydrogens:**
-
- ```python
- # Add explicit hydrogens
- mol_h = Chem.AddHs(mol)
-
- # Remove explicit hydrogens
- mol = Chem.RemoveHs(mol_h)
- ```
-
- **Kekulization and Aromaticity:**
-
- ```python
- # Convert aromatic bonds to alternating single/double
- Chem.Kekulize(mol)
-
- # Set aromaticity
- Chem.SetAromaticity(mol)
- ```
-
- **Replacing Substructures:**
-
- ```python
- # Replace substructure with another structure
- query = Chem.MolFromSmarts('c1ccccc1') # Benzene
- replacement = Chem.MolFromSmiles('C1CCCCC1') # Cyclohexane
-
- new_mol = Chem.ReplaceSubstructs(mol, query, replacement)[0]
- ```
-
- **Neutralizing Charges:**
-
- ```python
- # Remove formal charges by adding/removing hydrogens
- from rdkit.Chem.MolStandardize import rdMolStandardize
-
- # Using Uncharger
- uncharger = rdMolStandardize.Uncharger()
- mol_neutral = uncharger.uncharge(mol)
- ```
-
- ### 11. Working with Molecular Hashes and Standardization
-
- **Molecular Hashing:**
-
- ```python
- from rdkit.Chem import rdMolHash
-
- # Generate Murcko scaffold hash
- scaffold_hash = rdMolHash.MolHash(mol, rdMolHash.HashFunction.MurckoScaffold)
-
- # Canonical SMILES hash
- canonical_hash = rdMolHash.MolHash(mol, rdMolHash.HashFunction.CanonicalSmiles)
-
- # Regioisomer hash (ignores stereochemistry)
- regio_hash = rdMolHash.MolHash(mol, rdMolHash.HashFunction.Regioisomer)
- ```
-
- **Randomized SMILES:**
-
- ```python
- # Generate random SMILES representations (for data augmentation)
- from rdkit.Chem import MolToRandomSmilesVect
-
- random_smiles = MolToRandomSmilesVect(mol, numSmiles=10, randomSeed=42)
- ```
-
- ### 12. Pharmacophore and 3D Features
-
- **Pharmacophore Features:**
-
- ```python
- from rdkit.Chem import ChemicalFeatures
- from rdkit import RDConfig
- import os
-
- # Load feature factory
- fdef_path = os.path.join(RDConfig.RDDataDir, 'BaseFeatures.fdef')
- factory = ChemicalFeatures.BuildFeatureFactory(fdef_path)
-
- # Get pharmacophore features
- features = factory.GetFeaturesForMol(mol)
-
- for feat in features:
- print(feat.GetFamily(), feat.GetType(), feat.GetAtomIds())
- ```
-
- ## Common Workflows
-
- ### Drug-likeness Analysis
-
- ```python
- from rdkit import Chem
- from rdkit.Chem import Descriptors
-
- def analyze_druglikeness(smiles):
- mol = Chem.MolFromSmiles(smiles)
- if mol is None:
- return None
-
- # Calculate Lipinski descriptors
- results = {
- 'MW': Descriptors.MolWt(mol),
- 'LogP': Descriptors.MolLogP(mol),
- 'HBD': Descriptors.NumHDonors(mol),
- 'HBA': Descriptors.NumHAcceptors(mol),
- 'TPSA': Descriptors.TPSA(mol),
- 'RotBonds': Descriptors.NumRotatableBonds(mol)
- }
-
- # Check Lipinski's Rule of Five
- results['Lipinski'] = (
- results['MW'] <= 500 and
- results['LogP'] <= 5 and
- results['HBD'] <= 5 and
- results['HBA'] <= 10
- )
-
- return results
- ```
-
- ### Similarity Screening
-
- ```python
- from rdkit import Chem
- from rdkit.Chem import rdFingerprintGenerator
- from rdkit import DataStructs
-
- def similarity_screen(query_smiles, database_smiles, threshold=0.7):
- query_mol = Chem.MolFromSmiles(query_smiles)
- if query_mol is None:
- return []
-
- morgan_gen = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
- query_fp = morgan_gen.GetFingerprint(query_mol)
-
- hits = []
- for idx, smiles in enumerate(database_smiles):
- mol = Chem.MolFromSmiles(smiles)
- if mol:
- fp = morgan_gen.GetFingerprint(mol)
- sim = DataStructs.TanimotoSimilarity(query_fp, fp)
- if sim >= threshold:
- hits.append((idx, smiles, sim))
-
- return sorted(hits, key=lambda x: x[2], reverse=True)
- ```
-
- ### Substructure Filtering
-
- ```python
- from rdkit import Chem
-
- def filter_by_substructure(smiles_list, pattern_smarts):
- query = Chem.MolFromSmarts(pattern_smarts)
-
- hits = []
- for smiles in smiles_list:
- mol = Chem.MolFromSmiles(smiles)
- if mol and mol.HasSubstructMatch(query):
- hits.append(smiles)
-
- return hits
- ```
-
- ## Best Practices
-
- ### Error Handling
-
- Always check for `None` when parsing molecules:
-
- ```python
- mol = Chem.MolFromSmiles(smiles)
- if mol is None:
- print(f"Failed to parse: {smiles}")
- continue
- ```
-
- ### Performance Optimization
-
- **Use safe storage formats:**
-
- ```python
- import base64
- import json
- from pathlib import Path
- from rdkit import Chem
-
- # Portable exchange formats such as SMILES and SDF are safest for shared data.
- # For local caches, RDKit's binary molecule representation avoids generic pickle.
- payload = [base64.b64encode(mol.ToBinary()).decode("ascii") for mol in mols]
- Path("molecules.rdmol.json").write_text(json.dumps(payload))
-
- cached = json.loads(Path("molecules.rdmol.json").read_text())
- mols = [Chem.Mol(base64.b64decode(item)) for item in cached]
- ```
-
- Do not load Python pickle files from untrusted sources. Pickle deserialization can execute arbitrary code; prefer SMILES/SDF for interchange and RDKit binary payloads for trusted local caches.
-
- **Use bulk operations:**
-
- ```python
- from rdkit import DataStructs
- from rdkit.Chem import rdFingerprintGenerator
-
- # Calculate fingerprints for all molecules at once
- morgan_gen = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
- fps = [morgan_gen.GetFingerprint(mol) for mol in mols]
-
- # Use bulk similarity calculations
- similarities = DataStructs.BulkTanimotoSimilarity(fps[0], fps[1:])
- ```
-
- ### Version-Sensitive Behavior
-
- Pin RDKit versions when exact molecular identifiers or numeric features are part of a persisted dataset, model feature pipeline, or regulated report. Recent releases changed or documented behavior in several Python-facing areas:
-
- - **Canonical SMILES and stereo:** 2026.03 changed canonical double-bond handling to avoid stereo corruption, so some stereo-containing SMILES may differ from older releases.
- - **Descriptors and hashes:** 2024.09 corrected unbranched-alkane fragment descriptor SMARTS and changed some tautomer/protomer hash outputs.
- - **Drawing:** legacy `rdkit.Chem.Draw` canvas modules and functions such as `MolToImageFile`, `MolToMPL`, and `MolToQPixmap` were removed; use `Draw.MolToFile`, `Draw.MolToImage`, or `rdMolDraw2D`.
- - **MolStandardize:** use `rdkit.Chem.MolStandardize.rdMolStandardize`; the older Python MolStandardize implementation was removed.
- - **Similarity maps:** `GetSimilarityMapFromWeights()`, `GetSimilarityMapForFingerprint()`, and `GetSimilarityMapForModel()` now require an `rdMolDraw2D` drawing object.
-
- ### Thread Safety
-
- RDKit operations are generally thread-safe for:
- - Molecule I/O (SMILES, mol blocks)
- - Coordinate generation
- - Fingerprinting and descriptors
- - Substructure searching
- - Reactions
- - Drawing
-
- **Not thread-safe:** MolSuppliers when accessed concurrently.
-
- ### Memory Management
+ Twelve capability areas, each with worked code, are documented in
+ [references/core_capabilities.md](references/core_capabilities.md):
- For large datasets:
+ | # | Area | Covers |
+ | --- | --- | --- |
+ | 1 | Molecular I/O and creation | SMILES, MOL files and blocks, InChI, SDF and SMILES suppliers, multithreaded reading, writers |
+ | 2 | Sanitization and validation | disabling automatic sanitization, manual and partial sanitization, detecting problems first |
+ | 3 | Analysis and properties | atom and bond iteration, ring information and SSSR, chirality and stereochemistry, fragments |
+ | 4 | Descriptors | MW, LogP, TPSA, H-bond donors/acceptors, rotatable bonds, aromatic rings, bulk calculation, drug-likeness |
+ | 5 | Fingerprints and similarity | topological, Morgan/ECFP via `rdFingerprintGenerator`, MACCS, atom pair, torsion, Avalon; Tanimoto and other metrics; Butina clustering |
+ | 6 | Substructure searching | SMARTS queries, match retrieval, and a library of common patterns |
+ | 7 | Chemical reactions | reaction SMARTS, applying reactions, reaction fingerprints |
+ | 8 | 2D and 3D coordinates | depiction, template alignment, ETKDG embedding, force-field optimization, RMSD, constrained embedding |
+ | 9 | Visualization | single and grid images, substructure highlighting, custom drawer options, Jupyter integration, fingerprint bit environments |
+ | 10 | Molecular modification | explicit hydrogens, Kekulization, aromaticity, substructure replacement, charge neutralization |
+ | 11 | Hashes and standardization | Murcko scaffold and canonical hashes, regioisomer hashes, randomized SMILES for augmentation |
+ | 12 | Pharmacophore and 3D features | feature factories and feature extraction |
- ```python
- # Use ForwardSDMolSupplier to avoid loading entire file
- with open('large.sdf') as f:
- suppl = Chem.ForwardSDMolSupplier(f)
- for mol in suppl:
- # Process one molecule at a time
- pass
+ Worked workflows and the performance, thread-safety, and version-sensitivity notes are in
+ [references/workflows_and_best_practices.md](references/workflows_and_best_practices.md).
- # Use MultithreadedSDMolSupplier for parallel processing
- suppl = Chem.MultithreadedSDMolSupplier('large.sdf', numWriterThreads=4)
- ```
+ Prefer portable exchange formats (SMILES, SDF) for shared data; for local caches RDKit's
+ binary molecule representation avoids generic pickle.
## Common Pitfalls
1. **Forgetting to check for None:** Always validate molecules after parsing
2. **Sanitization failures:** Use `DetectChemistryProblems()` to debug
3. **Missing hydrogens:** Use `AddHs()` when calculating properties that depend on hydrogen
4. **2D vs 3D:** Generate appropriate coordinates before visualization or 3D analysis
5. **SMARTS matching rules:** Remember that unspecified properties match anything
6. **Thread safety with MolSuppliers:** Don't share supplier objects across threads
## Resources
### references/
This skill includes detailed API reference documentation:
- `api_reference.md` - Comprehensive listing of RDKit modules, functions, and classes organized by functionality
- `descriptors_reference.md` - Complete list of available molecular descriptors with descriptions
- `smarts_patterns.md` - Common SMARTS patterns for functional groups and structural features
Load these references when needing specific API details, parameter information, or pattern examples.
Only the files listed in `references/` and `scripts/` are bundled local resources. Names such as `rdkit`, `datamol`, `scipy`, and `sklearn` refer to installable Python packages, not local files in this skill.
### scripts/
Example scripts for common RDKit workflows:
- `molecular_properties.py` - Calculate comprehensive molecular properties and descriptors
- `similarity_search.py` - Perform fingerprint-based similarity screening
- `substructure_filter.py` - Filter molecules by substructure patterns
These scripts can be executed directly or used as templates for custom workflows.
-