Overview
Datamol Skill (K-Dense) is a procedural resource in the Scientific Agent Skills collection. It provides instructions and worked workflow references for using Datamol, a Python abstraction layer over RDKit, in cheminformatics and drug-discovery tasks. The Skill is not the Datamol package or an independently running agent: its role is to guide an assistant or researcher through molecular operations using separately installed software.
The documented workflow begins with molecular strings or structure files and proceeds through parsing, task-specific standardization, descriptor or fingerprint computation, and downstream analysis. Inputs include SMILES and file formats such as SDF and CSV; documented outputs include native rdkit.Chem.Mol objects, molecular representations, descriptor tables, similarity distances, clusters, scaffolds, conformers and visualizations. The Skill points to worked references for individual operations and broader load/filter/analyze, scaffold-series SAR and virtual-screening pipelines. It also explains how fingerprints can feed a separate machine-learning workflow rather than providing a predictive model itself.
Its practical guidance emphasizes preserving original structures and identifiers, recording rejected inputs, and choosing transformations appropriate to the scientific task. Standardization is not presented as guaranteed chemical repair. Environment changes can affect canonical representations and retained conformers, while full pairwise clustering can exceed available memory. Remote file workflows require provider-specific credentials, and the source explicitly excludes remote authorization and writes from its reported testing. These boundaries make the Skill useful as workflow guidance, not evidence that a particular dataset or scientific result has been validated.
Key Features
- Instructions for molecular parsing, format conversion, sanitization and task-specific standardization, including retention of failed source records.
- Worked workflow references covering descriptors, fingerprints, similarity distances, nearest-neighbour lookup, clustering and diversity selection.
- Scaffold and fragment analysis guidance, including Bemis-Murcko grouping and scaffold-disjoint train/test splits.
- Guidance for conformer generation, RMSD clustering, representative selection, SASA calculations and molecular visualization.
- File and batch-processing instructions covering molecular files, tabular outputs, parallel operations and credential-dependent remote I/O.
- End-to-end workflow references for library preparation and analysis, scaffold-series SAR and virtual screening.
Use Cases
- Suggested evaluation: prepare an external compound library while retaining original structures, identifiers, transformation records and rejected inputs.
- Suggested evaluation: explore a compound collection using fingerprints, similarity grouping and bounded diversity selection before selecting a screening subset.
- Suggested evaluation: organize an SAR series by scaffold and generate aligned molecular views for inspection.
- Suggested evaluation: construct scaffold-separated molecular features for a downstream model, checking that preprocessing and fingerprint settings remain consistent across splits.
How to Use
- Read the named Skill file to identify the relevant workflow and its environment requirements. Treat it as instructions, not a runnable agent.
- Prepare an isolated Datamol environment following the Skill’s setup section. Consult the linked Datamol documentation for upstream API details; keep the Skill and package roles separate.
- Supply molecular strings or structure files with stable source identifiers. Define a task-specific standardization policy before processing, and retain originals and parsing failures.
- Follow the Skill’s referenced workflow area for descriptors, similarity, scaffolds, conformers or file handling. Start with a small evaluation subset and inspect outputs before expanding the analysis.
- Record the environment and transformation choices. Check chemical invariants rather than expecting identical canonical strings or conformer counts across environments; budget memory for clustering and provide credentials only when remote I/O is needed.