Overview
RDKit Skill (K-Dense) is an individual resource in the Scientific Agent Skills collection. It supplies instructions, reference material, and optional helper scripts for applying RDKit to cheminformatics tasks; it is not the RDKit library itself or an independently running agent. The guidance favors RDKit when a workflow needs detailed molecular control, custom sanitization, or specialized algorithms, while identifying datamol as a simpler wrapper for standard workflows.
The documented workflow starts with molecular structures such as SMILES, SDF, MOL, or InChI and proceeds through parsing, validation, analysis, and optional transformation. Covered outputs include molecular descriptors, fingerprints, similarity results, substructure matches, reaction products, coordinates, and visualizations. Three bundled Python helpers address molecular properties, fingerprint-based similarity screening, and substructure filtering. Supporting references organize API details, descriptor definitions, SMARTS patterns, and workflow considerations.
Chemical interpretation remains an explicit responsibility. Canonical isomeric SMILES does not automatically reconcile tautomers, protonation states, salts, or unspecified stereochemistry. The helpers preserve source identifiers through invalid-record filtering and do not implicitly strip salts or neutralize structures. Fingerprint similarity is not an identity test, and the search helper accepts one valid query molecule rather than a query batch. Its historical pains option contains only five illustrative motifs, not the published PAINS catalogue. Descriptor thresholds, QED, and alerts are research heuristics rather than evidence of biological activity or safety. The helper scripts require an installed rdkit package; the Skill supplies guidance rather than that dependency.
Key Features
- Guidance for molecular input/output and validation, including SMILES, SDF, MOL, InChI, manual sanitization, and chemistry-problem detection.
- Descriptor and fingerprint workflows covering MW, LogP, TPSA, hydrogen-bond counts, Morgan/ECFP, MACCS, similarity metrics, and Butina clustering.
- SMARTS-based substructure searching and reaction SMARTS guidance, including match retrieval, reaction application, and reaction fingerprints.
- Instructions for 2D depiction, ETKDG embedding, force-field optimization, alignment, and checking failed conformer generation before further processing.
- Three bundled Python helpers: `molecular_properties.py`, `similarity_search.py`, and `substructure_filter.py`, usable directly or as workflow templates.
- Chemical-meaning safeguards covering original structure retention, source identifiers, enhanced stereochemistry in CXSMILES outputs, and explicit standardization policies.
Use Cases
- Suggested evaluation: build a descriptor table from a molecular collection while retaining original structures and source identifiers for traceability.
- Suggested evaluation: screen a compound collection against one query molecule using fingerprints, then inspect candidate matches without treating similarity as chemical identity.
- Suggested evaluation: filter structures with task-specific SMARTS patterns and verify the matches before interpreting alerts or excluding compounds.
- Suggested evaluation: prototype molecular depiction or conformer-generation workflows, checking parsing and embedding failures before downstream analysis.
How to Use
- Read the named Skill file to select the relevant workflow and distinguish its instructions from the RDKit dependency.
- Obtain the
skills/rdkitresources from the collection repository. Follow the Skill’s environment guidance for installingrdkit; avoid mixing conda and PyPI installations in one environment. - Prepare molecular inputs in a documented format. Preserve original structures and source IDs, validate parsed molecules, and decide explicitly whether any standardization is appropriate.
- Consult the reference files listed in the source instructions, then select a helper or adapt a worked workflow. Similarity search takes one valid query molecule.
- Evaluate a small, inspectable dataset first. Check invalid records, stereochemistry handling, SMARTS matches, and conformer failures; interpret descriptors and alerts as research heuristics, not validated biological conclusions.