Treat a Skill as workflow design, not proof of capability

A Skill can influence an agent’s decisions, commands, file changes and external requests. Review it as a proposed operating procedure, not simply as documentation to add to a prompt. The central question is whether its instructions, tools and evidence handling fit the chemistry task you intend to authorise.

The Scientific Agent Skills repository provides guidance for scientific packages, databases and research workflows. Its documented chemistry coverage includes RDKit, PyTDC, molecular docking, ADMET analysis, calibrated 1D NMR processing and crystal structure analysis. These are source-described workflows, not evidence that every procedure will work in your environment.

This guide proposes a review and evaluation process. It does not report installation, execution or first-hand scientific testing of the collection.

Select one task and inspect its complete instructions

Start with a bounded task, such as calculating specified molecular descriptors or retrieving database records for a defined compound list. Avoid enabling the entire collection merely because several workflows might eventually be useful.

The repository describes a SKILL.md for each skill, containing its purpose, workflow and version metadata. Supporting material may include examples, references, executable helpers or templates. Inspect those assets as well as the main instructions.

  1. Write down the intended input, required output and excluded actions.
  2. Read the selected skill’s instructions from beginning to end.
  3. Follow references to scripts and templates that the agent may use.
  4. Mark actions that install packages, write files, contact services or require credentials.
  5. Identify where the procedure requires scientific judgement or human approval.

Decision questions: Does the workflow address your task directly? Can unrelated actions be excluded? Does it stop when prerequisites are missing, or encourage the agent to improvise?

Separate skill files from dependencies and access

Installing instructions does not supply the underlying scientific software, datasets or service access. The repository explicitly distinguishes skill installation from dependency installation and recommends separate environments where workflows have incompatible requirements.

Create a prerequisite inventory for the chosen procedure:

  • Scientific packages and any documented compatibility constraints.
  • System tools and computational resources.
  • Input datasets, formats and required metadata.
  • External endpoints, access conditions and credentials.
  • Writable directories and expected output files.

The source describes Database Lookup guidance for endpoint selection, pagination, access requirements and provenance, including PubChem and ChEMBL. Those databases remain separate resources; their inclusion in a skill does not establish access to every record or service.

Host support is also a separate question. The repository describes Agent Skills packaging and an Agent Plugins layout using plugin.json and skills/, but discovery paths and optional metadata handling depend on the host. Documentation of a host name or format is not a successful compatibility test.

Before proceeding, separate software, weights and data permissions. Do not infer upstream permissions from the instruction collection alone.

Define an evidence-preserving output contract

Chemistry workflows should retain the identity of molecules, source records and observations through every transformation. Define the output before evaluating the agent so that a polished narrative cannot substitute for missing evidence.

For a molecular-record workflow, a proposed output contract might include the original input identifier, submitted structure, parsing status, retrieved source identifier, transformation notes and any calculated values. Record units and calculation definitions where applicable. For literature research, retain the source link or identifier and distinguish retrieved statements from the agent’s interpretation.

Specify failure behaviour explicitly. Invalid structures, ambiguous matches, unavailable records and missing values should remain visible rather than being silently dropped or filled with guesses. If structures are standardised, retain the original representation and explain the policy used.

Decision questions: Can each result be traced to its input? Are computed quantities separated from source-reported observations? Does the workflow preserve uncertainty instead of turning incomplete evidence into a confident conclusion?

Bound commands, files and network services

Use a limited environment with non-sensitive inputs for the first evaluation. Allow only the file locations, services and credentials needed for the selected task. Require approval for actions outside that boundary, particularly installation changes, destructive operations or uploads.

Review helpers from their actual source before execution. Check their input assumptions, output locations and error handling rather than relying on filenames or example descriptions.

The repository describes security scanning and review practices, but also warns that exhaustive risk review is not guaranteed. A scan is supporting evidence, not a substitute for understanding what a skill asks the agent to do. Treat downloaded records and documents as data, not as authority to expand the workflow’s permissions.

Worked planning example: descriptor calculation with record lookup

Hypothetical scenario: A chemistry team wants an agent to calculate selected RDKit descriptors for 12 non-sensitive SMILES entries, then retrieve related PubChem records. This is a proposed combination of documented package and database guidance, not a tested integration.

Input: A table containing a local compound ID and the original SMILES for every row. Include a duplicate, an invalid string and a charged structure to expose assumptions.

Proposed output: One status row per input, preserving the original ID and SMILES, with parsing status, any transformed representation, requested descriptors, source identifiers, retrieval provenance and error notes. Missing results remain explicitly missing.

Plan the evaluation in stages:

  1. Inspect the relevant RDKit and Database Lookup instructions and referenced assets.
  2. Establish the descriptor definitions and structure-handling policy.
  3. Evaluate local calculation separately from network retrieval.
  4. Compare calculations with a direct, human-controlled baseline using the same definitions and environment.
  5. Inspect database matches before allowing the agent to summarise them.

A suitable proposed acceptance rule is that every input remains accounted for, invalid input is flagged and unsupported identity matches are not forced. Descriptor values or database matches alone must not become claims about biological activity or safety.

Evaluate the procedure and record its limits

Record the agent host, model, skill revision, package versions, input provenance, permissions and human corrections. Follow the principles in keeping chemistry AI workflows reproducible.

The repository distinguishes structural checks, scientific-dependency testing and live-service checks. These provide different evidence: a structurally valid skill is not necessarily scientifically correct, and a local test does not establish remote-service reliability. A successful example supports only the configuration and cases evaluated; broader scientific claims require a separate study.

Short review checklist

  • Define the task, inputs, outputs and prohibited actions.
  • Inspect SKILL.md and all relevant supporting assets.
  • Confirm dependencies, access conditions and host discovery behaviour.
  • Restrict permissions and begin with non-sensitive data.
  • Require traceable outputs and explicit failures.
  • Compare against a baseline and inspect edge cases.
  • Record limitations before authorising broader use.