What a scientific Skill does—and does not do

A scientific Skill supplies procedural instructions for an assistant: which inputs to collect, which software to use, what checks to perform, and when to stop. Some resources also bundle helpers. Loading those instructions does not install dependencies, supply experimental data, or establish scientific validity. Host support for Skills differs, so instruction discovery and execution permissions need separate checks.

These ten resources cover three cheminformatics choices, one spectroscopy workflow, three materials resources, two thermodynamics/kinetics workflows, and one reactive-MD analysis resource. Their roles are not interchangeable: an ASE router prepares delegation, while a Cantera helper performs a narrowly defined calculation through separately installed software.

For the broader distinctions, read agents and Skills, MCP, Skills, agents and APIs, and the chemistry AI stack. The chemistry agent Skills review provides a complementary evaluation perspective.

For related guidance, see AI Skills for Chemistry · AI for Materials Science.

Compare the ten resources by role and data contract

In the table, K-Dense means the scientific-agent-skills collection; CCAS means computational-chemistry-agent-skills. Licence declarations concern the Skill or supplied collection code, not automatically its dependencies, databases or outputs.

Skill and collection Instruction/helper role; upstream software and dependencies Inputs → expected outputs Licence scope or unresolved status
RDKit Skill, K-Dense Detailed molecular guidance and three Python helpers; installed rdkit required SMILES/SDF/MOL/InChI → descriptors, fingerprints, matches, transformations, coordinates Skill declares BSD-3-Clause; collection declares MIT; scope unresolved
Datamol Skill, K-Dense Workflow instructions and references; Datamol wraps RDKit and returns native molecule objects Molecular strings/files → prepared molecules, descriptor tables, clusters, scaffolds, conformers Skill declares Apache-2.0; collection declares MIT; scope unresolved
rdkit-agent Skill, scottmreed/rdkit-agent Tool-use instructions for a separate RDKit WASM CLI, with documented MCP and Node.js interfaces Chemical notation/JSON collections → validation, conversion, descriptors, searches No supplied licence evidence
nmrglue Skill, K-Dense Instructions and calibrated 1D processing helper; nmrglue, NumPy and SciPy Complex FID plus acquisition/settings → spectrum.csv, report.json MIT for Skill/collection code; acquisition data separate
Pymatgen Skill, K-Dense Materials instructions and local helpers; pymatgen; optional mp-api client Structures, compositions, energy entries or query criteria → validation, conversion, hull analysis, query plans MIT for Skill; external software and data terms separate
ASE Skill, CCAS Top-level router only; Python and ASE, optional GPAW/MACE adapters Task intent/context → branch, rationale, missing inputs, next delegation Skill declares MIT; collection supplies LGPL v3 text; scope unresolved
Phonopy Skill, CCAS Displacement/analysis guidance; phonopy plus separate force provider Structure/settings and displaced-cell forces or force constants → assembly status, phonon-band/DOS/thermal files Skill declares LGPL-3.0-or-later; force providers separate
Cantera Skill, K-Dense Instructions and homogeneous-ignition helper; Cantera and NumPy Mechanism, initial state, reactor constraint, numerical settings → JSON, four histories, mechanism snapshot MIT for Skill/collection code; mechanism rights separate
pycalphad Skill, K-Dense Instructions and local equilibrium helper; pycalphad and NumPy TDB plus composition/phase/condition settings → report.json, phase-equilibria.csv MIT for Skill/collection code; TDB rights separate
ReacNetGenerator Skill, CCAS Reactive-MD tool-selection guidance; ReacNetGenerator and separate reacnet-md-tools wrapper; uv/python3 Trajectories, atom mapping, cell context → species/reaction artifacts, reports, logs Skill declares LGPL-3.0-or-later; packages and trajectories separate

Which cheminformatics route should you choose?

Choose Datamol guidance for routine molecular preparation, descriptors, scaffold grouping and diversity selection. Choose RDKit guidance when the task needs explicit sanitization decisions, detailed molecular manipulation or specialized algorithms. The distinction is interface and control, not evidence that either produces better scientific results.

Choose rdkit-agent guidance when structured CLI exchange fits the host. Its instructions emphasize checking overall_pass, examining repair suggestions, and bounding returned fields or matches. The underlying package—not the Skill—provides executable operations and MCP serving. Its documented MCP mode does not establish compatibility with every current host. Standard WASM stereoisomer enumeration is explicitly unsupported.

All three routes need a chemical-meaning policy. Preserve originals and identifiers before salt removal, neutralization or stereochemical changes. Canonical SMILES does not reconcile tautomers or establish pH-dependent protonation. Fingerprint similarity is not molecular identity, and descriptor cutoffs are not potency or safety predictions. RDKit's helper option named pains contains illustrative motifs rather than the published PAINS catalogue.

Proposed molecular-library evaluation

Proposed example, not a tested integration: evaluate Datamol for routine intake and RDKit for exception diagnosis using a small CSV with source_id and smiles. Include ethanol (CCO), benzene (c1ccccc1), a charged salt ([Na+].[Cl-]), unspecified stereochemistry and an intentionally incomplete ring string (C1CC).

  1. Intake: retain every source row. Produce accepted and rejected-record tables without silently deleting failures.
  2. Policy decision: decide whether salts and charges represent the intended entities. For this evaluation, retain them unless transformation is explicitly justified.
  3. Analysis: calculate selected descriptors and one recorded fingerprint schema. Export identifiers, transformed representations if any, settings and failure reasons.
  4. Inspection: verify that identifiers remain aligned after rejection; inspect any repair rather than accepting it as the original compound.
  5. Scale decision: assess memory before full pairwise clustering. The Datamol guidance warns that pairwise distances can exceed available memory; evaluate bounded diversity selection instead when appropriate.

A downstream model or DrugAgent handoff would be a proposed consumer of these artifacts, not established interoperability. Define its schema and provenance requirements first. If labels are added, evaluate scaffold/group separation before learned preprocessing; these Skills do not themselves supply a predictive model. See the RDKit AI workflow.

Keep spectroscopy processing separate from identification

The nmrglue helper targets uniformly sampled complex 1D FIDs. Its accepted inputs are a single complex fid array in .npz or a canonical complex time-domain NMRPipe file, with explicit acquisition parameters and manual phase settings.

The decision gate is acquisition evidence: spectral width, observation frequency, carrier position and complex sign convention must be established rather than guessed. Conflicting NMRPipe calibration metadata should stop intake. Experimental vendor decoding requires separate inspection; successfully opening a converted file does not validate the original decoding.

Inspect real and imaginary spectra before choosing signal-free baseline regions and integration windows. Outputs are positive peak candidates and signed integrals, not assignments or compound identities. Negative areas remain diagnostics. Zero filling improves interpolation, not acquired resolution; areas remain arbitrary signal-times-ppm quantities, not concentrations.

Proposed materials sequence: validation, routing, forces, analysis

A useful proposed evaluation combination is Pymatgen → ASE routing → separate force-provider workflow → Phonopy. This is an architectural suggestion, not a tested integration.

Start with Pymatgen guidance to inspect periodicity, coordinate mode, occupancies, warnings and symmetry sensitivity. Before conversion, identify representation loss and compare scientifically relevant properties after round-tripping.

Next, use the ASE top-level Skill to distinguish workflow preparation from calculator setup. Mixed requests route to ase/ase-workflows first, with ase/ase-calculators as a dependency. Its output is a delegation contract—not energies or forces. GPAW and MACE are named adapters, but the supplied router does not establish their detailed behavior. Requested execution is delegated through dpdisp-submit.

For Phonopy, specify cell context, supercell, displacement amplitude and analysis objective. A separate provider must calculate forces. Preserve displacement-to-force-file correspondence and stop if data are incomplete or units inconsistent. Phonon band paths, meshes and relevant long-range corrections require explicit decisions. Investigate imaginary modes against convergence and setup choices rather than automatically declaring instability.

Thermodynamics and kinetics answer different questions

Pymatgen local hull analysis consumes compatible computed-energy entries and supplied competing phases. pycalphad instead uses a thermodynamic TDB to calculate finite-temperature equilibrium phase fractions and compositions. Cantera's helper evolves a closed, adiabatic, homogeneous ideal-gas reactor using a kinetic mechanism. None substitutes for the others.

For pycalphad, supply exactly N-1 independent elemental mole fractions and a dependent non-vacancy element. Record selected and excluded phases, and inspect molar fractions, mass balance and sampling sensitivity. The bundled Cu-Ni database is hypothetical teaching material, not an assessed alloy database. Equilibrium does not predict precipitation rates or retained microstructures.

For Cantera, establish constant-volume versus constant-pressure conditions and preserve mechanism provenance. Its delay uses the global maximum temperature derivative on a uniform output grid. Insufficient heating or boundary maxima yield a null delay. Inspect refinement and conservation checks before comparison; numerical consistency does not validate the mechanism or experimental observable. These outputs can inform a proposed process study, but neither helper is a general process optimizer.

Reactive-MD analysis and failure recovery

ReacNetGenerator guidance separates routine trajectory analysis, lower-level control and inspection of existing outputs. It recommends rng-pipeline for routine LAMMPS dump workflows, native reacnetgenerator for lower-level options, and rng-query for existing species/reaction files.

Confirm atom-type mapping, scaled versus Cartesian coordinates and periodic-cell semantics before interpreting networks. The separate reacnet-md-tools wrapper addresses coordinate conversion, including triclinic cells. Ambiguous mappings require clarification, not guessed elements. Record whether HMM treatment was used; a quick first pass is not scientific validation.

Across workflows, distinguish three failures: invalid input, missing execution dependency and failed scientific/numerical checks. Preserve each status with logs and original artifacts. Do not replace missing forces, unresolved ignition delays or failed equilibria with plausible-looking results.

Licence boundaries and reader checklist

RDKit, Datamol and ASE have conflicting file-versus-collection licence declarations. Resolve the scope for the exact artifacts before reuse or redistribution; repository visibility alone establishes no licensing conclusion. Other listed Skill declarations also do not grant rights to force-provider software, kinetic mechanisms, thermodynamic databases or service data. Materials Project is an external data service accessed through mp-api, not an API supplied by the Pymatgen Skill.

Before evaluating a Skill:

  • Identify the scientific question and required output.
  • Separate instructions, helpers, wrappers, engines and external services.
  • Confirm host permissions and execution dependencies independently.
  • Record units, conventions, transformations and source identifiers.
  • Start with a bounded case containing deliberate failure inputs.
  • Inspect scientific assumptions as well as numerical checks.
  • Preserve settings, hashes, versions, warnings and unresolved decisions.
  • Check software and data licence scopes separately.

Sources

The primary instruction files support the roles and boundaries described here: K-Dense's RDKit, Datamol, nmrglue, Pymatgen, Cantera and pycalphad; CCAS's ASE router, Phonopy and ReacNetGenerator; and the rdkit-agent Skill. Compare the collection declarations in the K-Dense licence and CCAS licence when resolving scope.