RDKit’s place in an AI chemistry stack
An AI assistant can translate a research question into operations, but molecular calculations should come from an identifiable chemistry engine. RDKit fills that role: its official documentation describes molecular data structures, 2D and 3D operations, descriptor and fingerprint generation, and chemical searching. Its core algorithms are implemented in C++, with a Python interface and additional wrappers whose feature coverage should not be assumed identical.
RDKit is a toolkit, not a ready-made predictor of biological activity. A useful architecture separates the assistant’s planning from molecular processing, external information retrieval and interpretation. The assistant requests an operation; software executes it; a structured result records what happened; the assistant explains that result without inventing missing values.
This guide covers the complete structure-processing path. The PubChem agent integration guide is the companion for external retrieval, while MCP, Skills, agents and APIs explains the surrounding component roles.
Choose Python, MCP or Skill-guided execution
These routes are complementary, not interchangeable. Choose according to who controls execution, which chemistry operations are required and how much auditability the application needs.
| Resource or route | Actual role | When to consider it | Boundary to check |
|---|---|---|---|
| RDKit Python interface | Direct access to toolkit operations | Explicit batch pipelines, custom sanitization and controlled parameters | Define your own record schema, failure handling and environment |
| RDKit MCP Server (TandemAI) | Python server exposing RDKit tools to an MCP-capable client | Assistant-driven tool calls | Inspect the actual tool inventory; comprehensive coverage is a stated goal, not demonstrated coverage |
| RDKit Skill (K-Dense) | Instructions, references and optional Python helpers | Teaching an assistant detailed RDKit workflows | The rdkit dependency must be installed separately |
| rdkit-agent Skill (scottmreed) | Instructions for a separate RDKit WASM-powered CLI | Validation-first, structured exchanges through the documented CLI | WASM functionality differs from Python; stereoisomer enumeration is documented as unsupported |
| Datamol Skill (K-Dense) | Guidance for using Datamol, a Python abstraction over RDKit | Routine preparation, batch analysis and scaffold workflows | Defaults still need a task-specific chemical policy |
| PubChem PUG REST | Hosted HTTP API for selected chemical information | Identifier, record and property enrichment | Retrieval is not local chemistry execution or experimental validation |
Loading a Skill does not install its execution dependencies. Hosts also differ in how they load instructions and permit tool use. For the TandemAI server, inspect exposed tools and settings before selecting a client; dependency declarations and a current protocol overview do not establish tested interoperability. The chemistry MCP selection guide and AI Skills for chemistry guide help frame those decisions.
For connection and tool discovery, see MCP for chemistry.
Define inputs and an auditable record contract
Start with a scientific question: are you comparing submitted substances, parent molecular graphs or a deliberately standardized representation? That decision determines what information can be transformed and what must remain untouched.
The K-Dense RDKit Skill documents SMILES, SDF, MOL and InChI workflows. Its Datamol counterpart also discusses tabular inputs such as CSV. These are documented workflow options, not proof that every MCP tool accepts each format.
A proposed record contract should include:
- Source: stable local record ID, source identifier, input format and original structure or file reference.
- Processing: parsing status, validation findings, standardization policy ID and transformation history.
- Chemistry: retained representation, salt/component treatment, charge treatment and stereochemistry status.
- Results: requested descriptors, fingerprint specification, search settings and returned matches.
- Provenance: software environment, execution route, external identifiers and retrieval date where applicable.
- Failures: stage, diagnostic, attempted action and final disposition.
Keep source IDs aligned even when records fail. A compact accepted-molecule list without a rejection ledger can break the relationship between features and source rows. The dataset provenance guide provides a broader framework for maintaining that relationship.
Parse first, then standardize deliberately
Parsing and sanitization decide whether a structure can enter the requested operation. Standardization selects a representation for a particular task. Neither establishes that the submitted structure is the intended compound.
The K-Dense RDKit guidance explicitly distinguishes canonical isomeric SMILES from reconciling tautomers, choosing protonation states, removing salts or resolving unspecified stereochemistry. Its helpers are described as preserving source IDs and not implicitly stripping salts or neutralizing molecules.
Before processing, make three decisions:
- Components and salts: retain the complete submitted structure, create a separate parent representation, or exclude unsuitable multicomponent records? Record any removed component.
- Charges: retain formal charges or apply a justified transformation? Adding hydrogens does not select protonation at a specified pH.
- Stereochemistry: preserve specified information and mark unspecified centers rather than filling them in silently. Where enhanced stereochemistry matters, choose an exchange representation that retains it.
Use direct RDKit when custom sanitization or specialized algorithms matter. Datamol guidance can simplify routine operations, but metal disconnection, neutralization and salt removal can change the entity being studied. These transformations should not be automatic for formulation or organometallic tasks.
Compute features and distinguish search questions
After acceptance under the chosen policy, descriptors provide numerical features. The supplied RDKit Skill covers molecular weight, LogP, TPSA and hydrogen-bond counts, among other properties. Store descriptor names and computation settings alongside values. Do not replace a failed calculation with an invented number or silently treat missing output as zero.
Fingerprints encode selected molecular features for comparison or downstream modelling. Record the family, parameters, size where applicable and chirality setting. The Skill describes Morgan/ECFP and MACCS workflows, but the appropriate choice remains a task-specific evaluation.
Similarity search asks which representations are close under a specified fingerprint and metric. It is not an identity test: the RDKit Skill warns about fingerprint collisions and chirality settings. Substructure search asks whether a SMARTS query matches a molecular graph under particular matching rules. Save the query, settings, match counts and result limits so the answer remains interpretable.
For database workflows, RDKit’s PostgreSQL cartridge supports similarity and substructure searches plus descriptor calculations. That is a separate deployment route, not evidence that an MCP wrapper exposes the cartridge. Fingerprints can also feed a separate predictive model; they do not supply labels or demonstrate predictive validity by themselves.
Proposed example: a traceable aspirin-neighbour search
The following is a proposed evaluation workflow, not a report of executed calculations. Its goal is to retrieve a reference structure, prepare a small local collection and produce a reviewable similarity shortlist.
Inputs: a local table with stable row IDs and submitted SMILES, plus PubChem CID 2244 as the reference identifier. Include malformed, charged, multicomponent and stereochemically unspecified records deliberately to exercise failure handling.
- Retrieve the documented CID 2244 SDF example. Preserve the response, identifier and retrieval date; validate the returned structure locally.
- Parse local inputs while retaining every original row. Quarantine malformed or empty structures rather than guessing their identity.
- Apply a declared example policy: retain the full submitted graph, retain formal charges, preserve specified stereochemistry and leave unspecified stereochemistry unresolved. Do not strip salts or select tautomers in this initial pass.
- Calculate a selected descriptor set and a consistently configured Morgan fingerprint for accepted records and the reference. Explicitly enable chirality if it belongs in the comparison.
- Rank accepted records using the chosen similarity metric, returning a bounded shortlist. Separately run a reviewed SMARTS query if the question requires a particular motif; do not confuse motif membership with fingerprint ranking.
- Review multicomponent records before interpreting the shortlist. If parent-only comparison is later justified, create a separately labelled branch rather than overwriting the first result.
Outputs: accepted-record features, ranked neighbours, substructure results where requested, a transformation ledger and rejected records. Decision points: whether the reference parsed successfully, whether each record satisfies the policy, and whether the selected representation answers the scientific question. No similarity score or descriptor result is asserted here.
Add agent calls and bounded PubChem enrichment
A recommended architecture is:
User question → assistant plan → structured validation request → chemistry execution → result checks → optional PubChem enrichment → provenance-backed explanation.
This is a proposed integration pattern, not tested interoperability. Keep the chemistry layer responsible for calculations and the assistant responsible for selecting approved operations and explaining returned evidence.
The rdkit-agent Skill recommends inspecting overall_pass, corrected_values and fix_suggestions, using JSON exchange and bounding returned fields or matches. Treat repair suggestions as candidate changes requiring inspection, not authorization to replace the original structure. For TandemAI, inspect tool schemas rather than assuming the whole RDKit API is available.
PubChem PUG REST supports compound identifiers, selected properties and structure-oriented searches. For example, the documented formula and InChIKey request can enrich the reference record. Keep retrieved properties distinct from local calculations and investigate discrepancies rather than silently choosing one.
Respect published request limits and dynamic throttling, with bounded retries. The tutorial positions PUG REST for focused retrieval, not millions of individual calls; bulk collection needs an appropriate bulk-download workflow.
Handle failures and scientific limitations
Separate invalid chemistry, unsupported operations, malformed tool requests and service failures. Route each to an actionable outcome: quarantine, correct the request, use an explicitly selected alternative engine, or defer retrieval. Do not repeatedly retry a chemically invalid input.
The standard WASM build described by the rdkit-agent Skill returns NOT_SUPPORTED_IN_WASM for stereoisomer enumeration. If enumeration is required, evaluate a Python route separately. Optional 3D stages also need failure gates: the RDKit guidance says failed embedding must not proceed into force-field optimization. For clustering, the Datamol Skill warns that full pairwise distances can exceed memory; evaluate bounded alternatives before scaling.
Descriptor cutoffs, QED and structural alerts are research heuristics, not evidence of potency, safety or synthetic feasibility. The RDKit helper’s historical pains option contains only illustrative motifs, not the complete published PAINS catalogue. TandemAI’s evaluation suite and LLMJudge can help assess agent behaviour, but their existence does not establish scientific accuracy.
Before redistributing Skill materials, resolve their licensing scope: the K-Dense RDKit and Datamol files declare different licences from the collection-level MIT licence. Those unresolved declarations do not affect what a calculation means, but they matter when packaging instructions or helpers. Source visibility alone is not a licence grant.
Reader checklist
- Define the scientific entity and preserve original structures and IDs.
- Confirm installed dependencies, exposed tools and required input schemas.
- Record salt, charge, tautomer and stereochemistry decisions explicitly.
- Keep failures aligned with source rows; inspect repairs before acceptance.
- Save descriptor, fingerprint and search settings with outputs.
- Separate local calculations from retrieved PubChem values.
- Evaluate a small challenge set before scaling or adding a predictive model.
- Check unsupported features, memory limits, throttling and redistribution terms.
Continue with these related guides: APIs for Chemistry AI Applications: Getting Started with PubChem and RCSB PDB.
Sources
- RDKit overview and official README: toolkit operations, interfaces and database searching.
- TandemAI README: MCP tools, client, tool discovery and evaluations.
- K-Dense RDKit Skill, Datamol Skill and rdkit-agent Skill: documented workflows and limitations.
- PUG REST specification, tutorial and programmatic-access overview: retrieval operations and service boundaries.