Start with the scientific decision

Build the smallest workflow that can support a defined research decision. A database lookup, molecular property screen and crystal relaxation need different inputs, evidence and stopping rules. Do not start by connecting every available agent component.

Write a specification covering the chemical domain, supplied structures or records, desired output, units, unacceptable errors and approving researcher. State whether the result is a retrieved observation, calculated descriptor, model estimate or proposed experiment. For reaction planning, a suggested route is not evidence that the reaction will succeed.

The architecture below is a planning framework. Resource capabilities are source-described; connections between separately listed projects are proposed unless explicitly documented. No combined scientific testing is claimed.

Related guides: Materials Science Workflow.

Separate the stack into responsibilities

An LLM should interpret requests and explain evidence; numerical and structural work should be assigned to explicit tools. The following role table helps distinguish components that are often grouped together as “AI.”

Layer Responsibility and example Boundary to preserve
LLM Proposed request interpretation, tool selection and narrative synthesis Its answer is not a chemical measurement
Agent ChemGraph orchestrates construction, simulation, analysis and reporting; ChemCrow connects questions to chemistry tools Execution depends on tools, settings and external dependencies
Skill Scientific Agent Skills packages procedures and supporting assets Instructions are not an installed scientific runtime
MCP server RDKit MCP Server (TandemAI) exposes RDKit functions to clients A wrapper does not replace RDKit or establish complete coverage
API and client PubChem PUG REST supplies HTTP retrieval; PubChemPy supplies a Python interface Both access upstream PubChem, not a new database
Scientific library RDKit provides molecular operations, descriptors and fingerprints Feature generation is not endpoint-specific prediction
Predictor toolkit Chemprop supports message-passing model training, evaluation and prediction A usable predictor needs a selected model and data contract
Data and evaluation Therapeutics Data Commons supplies datasets, splits, metrics and benchmarks Dataset provenance and terms need their own checks
Integrated platform Makya offers vendor-described molecular generation and candidate assessment Its external-model API does not establish integration with this stack

A project can span several roles. Choose by the responsibility needed, not by its label. The MCP, Skills, agents and APIs guide provides a companion vocabulary for these distinctions.

Choose direct calls, Skills or MCP deliberately

Use direct library or API calls when the task is fixed and its parameters are already known. Add an agent when the workflow needs clarification, branching or coordinated tool use.

A Skill can guide direct code execution without MCP. The RDKit Skill (K-Dense) documents parsing, validation, descriptors, similarity and structure manipulation, with optional helpers. It still requires RDKit. The Pymatgen Skill (K-Dense) guides structure validation, symmetry sensitivity, conversion planning and local phase diagrams; its Materials Project helper defaults to offline query planning.

MCP instead provides a client-server interface for tools and context. The protocol distinguishes host, client and server, with tools, resources and prompts as separate primitives. Check the actual host and server capabilities: the current protocol overview does not establish compatibility for older project implementations.

For example, PubChem MCP (cyanheads) offers read-only tools and URI-templated resources, but its README notes that some clients expose tools only. TandemAI supplies a tool-listing utility; inspect that inventory rather than treating its ambition to expose every RDKit function as achieved coverage.

Skill discovery and optional metadata behavior vary by host. Loading instructions does not install packages, calculators or credentials. Consult chemistry MCP selection when deciding whether another server adds value.

Sketch an architecture with explicit handoffs

A proposed architecture is:

Research request → LLM and agent host → selected tools → retained artifacts → scientific review.

Skills provide procedural guidance alongside this flow. Tool branches can call libraries directly, retrieve data through an API, or invoke MCP servers. There is no mandatory serial chain through every layer.

For ChemGraph, single_agent is the documented starting workflow. Its CLI and asynchronous Python interface can connect requests to chemistry tools and session artifacts. More elaborate delegation is optional, not a prerequisite for a molecular lookup.

Define a contract at every handoff:

  • Input: original record ID, representation, identifier namespace, units, requested operation and approved scope.
  • Output: source IDs, raw artifacts, derived values, method or model identity, warnings and explicit completeness status.
  • Decision: proceed, request clarification, quarantine a record, retry within bounds or stop.

Do not assume RDKit features match a Chemprop configuration. Verify the selected tutorial and model requirements. For periodic structures, record lattice, coordinate convention, occupancies and boundary conditions before conversion. The chemistry APIs guide helps frame retrieval contracts separately from agent behavior.

Proposed molecular example: a property prioritization screen

Suppose a team wants to rank a small candidate set by aqueous solubility under a specified measurement context. This is a proposed evaluation, not a tested integration or a prediction result.

  1. Specify the endpoint. Input candidate SMILES with source IDs and labeled reference observations. Define units, conditions and the intended chemical domain. If reference measurements do not match the endpoint, stop model development rather than substitute a convenient label.
  2. Resolve identities. Propose PUG REST or PubChemPy for retrieval, or PubChem MCP for an assistant host. Preserve all candidate matches and CIDs. The documented PUG REST property request for aspirin, CID 2244, offers a non-sensitive retrieval check; it is not a solubility training example.
  3. Prepare structures. Evaluate direct RDKit processing, guided by the RDKit Skill, or use TandemAI after inspecting its exposed schemas. Keep originals and log invalid inputs. Make salt, tautomer, protonation and stereochemistry policies explicit; canonical SMILES alone does not settle them.
  4. Develop a predictor. Investigate whether a TDC dataset matches the endpoint, then evaluate an appropriate Chemprop workflow. Preserve split assignments, transformations, model settings and exclusions. TDC documents scaffold splitting, but the split must answer the intended generalization question.
  5. Evaluate before ranking. Compare held-out estimates with reference labels using endpoint-appropriate metrics. Inspect chemical-domain differences and excluded records before accepting a ranking.

The proposed output table should contain source ID, resolved structure, prediction, units, model reference and eligibility status. Retrieved properties and computed descriptors must remain distinguishable from predictions. Use the RDKit AI workflow guide for focused molecular preparation.

Makya is a separate SaaS design option if candidate generation is needed. Its product page describes multi-objective generation and API connections, but export schemas, authentication and an integration with this example remain unspecified.

Proposed materials example: retrieve, validate, then plan relaxation

Suppose a team wants to inspect candidate silicon structures and prepare selected records for relaxation. This proposed workflow separates database evidence from a new calculation.

Retrieve: Materials Project MCP (benedictdebrah) documents element, band-gap and stability searches through mp_api, plus structure retrieval by material ID. Plan bounded filters and requested fields, then approve network access. An API key is required. Missing properties should remain missing, not become zero-valued filter inputs.

Validate: Propose the Pymatgen Skill for retaining parser warnings, checking periodicity and occupancies, and examining symmetry across tolerances. Output a validation report and provenance manifest before approving conversion. Local hull analysis requires compatible energies and competing phases; computed hull membership is not experimental stability.

Prepare execution: The ASE Skill (Jinzhe Zeng Group) is a router, not a calculator. It selects ase/ase-workflows or ase/ase-calculators, reports missing inputs and delegates the next step. The top-level Skill does not execute calculations; requested execution is delegated through dpdisp-submit.

ChemGraph is another orchestration option for ASE calculations, with available engines detected from the environment. Any handoff from these separate Skills or the Materials Project server to ChemGraph is proposed. Review calculator suitability, units, constraints and convergence criteria before submission. EMT is described as a setup-check calculator, not general high-accuracy chemistry.

The desired output is a reviewed job plan, followed—only after approval—by calculation artifacts and a convergence assessment. See chemistry and materials AI Skills for procedural selection.

Handle failures without changing the scientific question

Propose separate checks for transport success, schema validity, chemical meaning and scientific adequacy. Passing one does not imply passing the next.

For retrieval, use rate limiting and bounded retries consistent with service documentation. PUG REST is intended for focused requests, not unrestricted bulk per-record access. Preserve partial responses and pagination state. PubChem MCP distinguishes missing records, absent GHS data and unavailable 3D structures; none should be interpreted as a negative scientific finding.

For local processing, quarantine invalid structures and preserve their IDs. Failed embedding should not proceed into optimization. For simulation, treat parser success separately from convergence, and review results before comparison. For prediction, evaluate whether held-out evidence supports the intended decision rather than merely checking that a table was produced.

TandemAI’s evaluation suite includes LLM judging, but judge approval is not independent chemical validation. ChemCrow’s public package excludes tools described in its paper and explicitly cannot reproduce the same results. Paper results therefore cannot be assigned to this installation.

Plan permissions, deployment and reproducibility

Begin in an isolated environment with non-sensitive examples and a restricted tool set. Approve external queries, uploads, file mutations and calculation submission separately. Keep credentials out of prompts and saved reports; treat retrieved documents as evidence, not instructions.

Local execution does not automatically mean local-only data handling: hosted LLMs and databases introduce external transfers. Containers also do not make broad filesystem access safe by themselves. ChemGraph specifically warns that workspace shell access is not confined to the workspace. Its Kubernetes templates require review of authentication and persistence before deployment.

Record package and Skill revisions, model identifiers, calculator settings, query filters, retrieval time, raw responses, checksums, transformations and failures. Save machine-readable outputs alongside narratives. Keep incompatible dependencies in separate environments and re-evaluate after updates; see reproducible chemistry AI.

Check software, weights, data and service terms separately. Two Skill scope conflicts remain unresolved: the RDKit Skill declares BSD-3-Clause while its repository declares MIT; the ASE router declares MIT while its repository license contains LGPLv3 text. Clarify scope before redistribution rather than assigning a single blanket license.

Before extending autonomy:

  • Is the decision and endpoint explicit?
  • Are input/output contracts and stopping rules recorded?
  • Are identity ambiguity and missing data handled?
  • Are dependencies, host support and permissions checked?
  • Can every value be traced to a source or computation?
  • Has scientific evaluation—not just interface testing—been planned?

Sources

Primary references for the architecture and workflow boundaries include the MCP overview, Agent Skills format overview, ChemGraph README, TandemAI README and PUG REST specification.

Scientific procedures are described in the named RDKit Skill, Pymatgen Skill, ASE router and Materials Project MCP README.

Supporting roles and limitations come from the RDKit overview, Chemprop documentation, TDC README, PubChemPy README, PubChem MCP README, ChemCrow README, Scientific Agent Skills README and Makya product page.