Compare scientific roles, not agent labels
Choosing a chemistry agent starts with the artifact you need: a simulation trajectory, a MOF structure, a worked chemical-reasoning solution, a hypothesis, or an implemented machine-learning experiment. These are different deliverables, so a single ranking would obscure more than it explains.
The eight projects below also differ in evidence depth. Some provide software and setup guidance; others expose research code with incomplete deployment instructions. DrugAgent is described in a paper that links code, but the supplied repository page does not establish usable implementation contents. Repository visibility alone is not proof of an applicable open-source licence.
This guide compares source-described capabilities and then proposes evaluation steps. It does not report first-hand execution or scientific testing. For broader orientation, see AI agents for chemistry; for architectural distinctions, use agents and Skills.
Classify the systems by what they coordinate
Four useful categories emerge from the sources:
- Calculation orchestration: ChemGraph connects requests to molecular construction, calculators, analysis and reporting; MDCrow concentrates on molecular dynamics.
- Domain-specific materials tools: ChatMOF routes MOF retrieval, prediction and generation tasks through dedicated modules.
- Reasoning and proposal development: SciAgents generates graph-guided hypotheses; ChemAgent (Gerstein Lab) investigates memory-supported chemical reasoning.
- Tool learning and computational research: ChemAgent (AI4Chem) studies planning and tool-use training; STELLA coordinates biomedical research tools; DrugAgent develops drug-discovery ML implementations.
These are frameworks, not interchangeable standalone predictive models. Their scientific outputs depend on the underlying engines, models, databases and task definitions.
Skills and templates are another layer: ChemGraph Skills supply instructions, while STELLA templates encode reusable workflows. Neither should be counted as an additional autonomous scientist. Loading instructions does not install execution dependencies, and host support differs. A hosted interface likewise does not establish that every local capability is available online. The chemistry AI stack explains how these layers fit together.
Eight agents: inputs, outputs and tool scope
This comparison uses documentation retrieved on 2026-10-06 UTC. The table describes roles, not comparative accuracy. Follow each resource link for project-specific context.
| Resource | Scientific role and input | Source-described output | Tools and dependencies that matter |
|---|---|---|---|
| ChemGraph | Chemistry requests and calculation settings; specialized workflows accept documents or ligand/receptor information | Structures, calculation results, trajectories, spectra and reports | LangGraph, ASE, RDKit and MCP; available calculators depend on installed engines |
| MDCrow | Natural-language MD tasks, including a protein identifier and requested conditions | Simulation execution and requested analysis, such as RMSD over time; complete file formats are unspecified | LangChain orchestration, OpenMM simulation, a configured environment and an LLM-provider key |
| ChatMOF | Textual MOF queries and desired properties | Retrieved information, property predictions or generated structures | Task-module setup; generation additionally requires GRIDAY; online demonstration is search-focused |
| SciAgents | Keywords or sampled knowledge-graph paths | Structured hypothesis JSON and expanded proposals with critique | GraphReasoning, graph and embedding files, OpenAI and Semantic Scholar APIs; automated workflow uses AG2 |
| ChemAgent (Gerstein Lab) | Multi-step chemical-reasoning problems | Memory-supported solutions and stored task experiences | Memory construction and model configuration; complete installation guidance is missing |
| ChemAgent (AI4Chem) | Chemistry/materials tasks and tool-call parameters | Tool-supported answers or predictions; concrete schemas are unspecified | ChemToolBench and the chemistry/materials tool pool; deployment requirements are not established |
| STELLA | Biomedical objectives; case studies also use prior screening feedback | Evidence-backed gene rankings or enzyme-variant proposals; case scripts save CSV and text | OpenRouter key locally; optional literature/search services and separate biomedical resources |
| DrugAgent (Liu et al.) | Task description, starter files and evaluator | Implemented ML solutions, performance or failure reports, and experimental submission files | Planner/Instructor workflow with domain documentation; accessible implementation and setup remain unconfirmed |
ChemGraph explicitly documents serving and consuming MCP tools. That is not evidence that all eight projects support the same protocol interfaces or current clients. Proposed cross-project connections require interface inspection and separate testing.
Code, licences and reproduction barriers
Record the paper snapshot separately from the repository revision. A current README can describe functionality absent from the original paper, and a paper’s code-availability statement does not establish a complete runtime.
| Project | Paper or research snapshot supported here | Code/licence evidence | Main reproduction barrier |
|---|---|---|---|
| ChemGraph | README cites a Communications Chemistry article, volume 9, article 33 (2026) | Repository code and documented entry points; Apache-2.0 | Matching calculator dependencies and workflow-specific infrastructure |
| MDCrow | arXiv:2502.09565v1 | Repository code and installation guidance; MIT | OpenMM environment, provider integration and task settings |
| ChatMOF | README cites Nature Communications 15, 4705 (2024) | Package configuration and CLI; MIT | Local prediction/generation modules and GRIDAY, not just demo access |
| SciAgents | README cites Advanced Materials, article 2413523 (2024) | Notebook code; Apache-2.0 licence text conflicts with an MIT package classifier | External graph assets, APIs and unresolved licence metadata |
| ChemAgent (Gerstein Lab) | arXiv:2501.06590v1 | Released research code; Apache-2.0 badge alone does not confirm licence text | Placeholder installation section and prerequisite memory construction |
| ChemAgent (AI4Chem) | CheMatAgent, arXiv:2506.07551v2 | Repository framework and dataset locations described; licence unknown | Missing deployment instructions and unconfirmed checkpoints |
| STELLA | arXiv:2507.02004v1; current README also links an expanded bioRxiv presentation | Code and entry points described; Apache-2.0 | Separately distributed resources and changing literature/model outputs |
| DrugAgent | arXiv:2411.15692v2 | Paper links an anonymous repository; supplied page exposes no implementation files or code licence | Establishing code access, environment and evaluator reproduction |
SciAgents’ licence conflict matters before redistribution or adaptation: inspect the intended code snapshot and seek clarification rather than selecting whichever label is convenient. STELLA’s README and package metadata also have differing version labels, so do not use either as a unified release identifier.
Code licences do not automatically cover model weights, datasets, APIs or hosted access. ChemGraph’s README, for example, separately identifies ASL terms for MACE-Polar weights. See software, weights and data licences for this distinction.
Distinguish the two ChemAgent projects
ChemAgent (Gerstein Lab) studies reusable memory for chemical problem solving. Task splitting, execution, association and reflection work with Plan Memory, Execution Memory and Knowledge Memory. Its documented experiments use SciBench-derived atkins, chemmc, matter and quan datasets. Drug discovery and materials science are future possibilities in the paper, not established end-to-end workflows.
Reproduction should begin with memory provenance and construction. The README alternates between image and img for Knowledge Memory and between Xagent and XAgent in paths; resolve these against the selected code before running experiments. Refinement after failure and evaluation enablement are separate controls.
ChemAgent (AI4Chem) is the framework in ChemistryAgent associated with the paper named CheMatAgent. Its research emphasis is external-tool selection, parameter filling and execution. HE-MCTS separates planning optimization from execution, while ChemToolBench supports training and evaluation. This is not Gerstein Lab’s memory system, nor does the brief README establish a ready-to-run deployment.
Use the repository and paper identifier together in notes, dependency records and citations; the shared name is insufficient.
Choose by research stage and intended use
For computational chemistry, start by deciding whether you need general molecule/calculator orchestration or a specialized MD workflow. ChemGraph and MDCrow are relevant candidates respectively; neither substitutes for method selection.
For materials discovery, ChatMOF fits MOF-specific retrieval, prediction and generation, whereas SciAgents fits hypothesis development for bio-inspired materials. ChemAgent (AI4Chem) is relevant to studying chemistry/materials tool learning, with a less established deployment path.
For literature research, STELLA offers biomedical searches and evidence-oriented workflows; SciAgents uses retrieval to critique hypotheses. Retrieval scope is not a guarantee of comprehensive coverage.
For drug discovery, distinguish target or variant prioritization from ML programming. STELLA describes biomedical prioritization workflows; DrugAgent studies implementation and evaluation on PAMPA, HIV and DAVIS binary-classification case studies. Its authors explicitly caution against direct pipeline deployment. AI for drug discovery provides broader workflow context.
For learning, inspect architecture and small documented examples. For reproduction experiments, choose a fixed task and preserve artifacts. For actual research, require method review, evidence traceability and independent checks before advancing outputs. These are proposed suitability criteria, not readiness certifications.
For stage-specific materials choices, see AI for Materials Science.
Proposed example: evaluate an MD assistant without overclaiming
A concrete proposed pilot is to evaluate MDCrow using the README’s task involving protein 1ZNI, 300 K, a 0.1 ps simulation and RMSD over time. This is a workflow check, not a scientifically sufficient sampling study.
- Define inputs: Record the protein identifier, intended structure, temperature, duration and RMSD request. Have a researcher specify acceptable preparation, force-field and analysis choices before execution.
- Check prerequisites: Follow the documented environment and provider setup. Confirm that the required simulation tools are available rather than assuming the language interface supplies them.
- Inspect decisions: Examine structure preparation, simulation configuration, atom selection and alignment choices before accepting execution.
- Collect outputs: Request available configuration records, execution logs and RMSD results. Treat absent artifacts as a pilot failure; their formats are not fully established by the supplied README.
- Decide whether to continue: Compare the recorded setup and analysis with a researcher-prepared reference workflow. Separate successful execution from scientifically meaningful dynamics.
A proposed downstream DrugAgent study could use a separately curated dataset, starter files and evaluator to investigate ML implementation. No automatic MDCrow-to-DrugAgent handoff or tested interoperability is established. First confirm code access and define a data contract. Reproducible chemistry AI covers the recordkeeping needed for such evaluations.
Handle failures and scientific limitations
Classify failures before retrying. Missing engines, inaccessible assets or provider credentials are environment failures; malformed tool parameters are orchestration failures; unconverged calculations or unsuitable methods are scientific failures. Preserve the failing inputs and logs, then correct the relevant layer. Do not let an agent silently substitute a different method to obtain a successful-looking answer.
Apply project-specific gates. ChemGraph’s EMT examples are setup checks, not general high-accuracy chemistry. Its workspace shell is not confined to the workspace despite action reviews. ChatMOF’s online search demo cannot stand in for local prediction or generation. SciAgents’ proposed mechanisms remain hypotheses. STELLA warns that gene rankings can vary with model stochasticity and literature-index updates. DrugAgent performance reports must be checked against actual evaluator artifacts, especially given the paper’s concern about fabricated results.
For comparison experiments, use equivalent inputs, documented methods, fixed data snapshots and explicit failure categories. Do not combine reasoning benchmarks, retrieval tasks and simulation completion into one score. More agents, more tools or longer proposals do not establish better scientific conclusions.
Reader checklist
- Identify the required artifact and research stage.
- Record both repository revision and paper version.
- Confirm code contents, installation guidance and required assets.
- Resolve licence uncertainty separately for code, weights and data.
- Specify inputs, units, method choices and acceptance criteria.
- Preserve tool calls, configurations, logs and raw outputs.
- Set stopping conditions for missing evidence or scientific failure.
- Review proposed integrations before attempting a handoff.
Sources
Primary capability sources are the official READMEs for ChemGraph, MDCrow, ChatMOF, SciAgents and STELLA.
The ChemAgent distinctions draw on the Gerstein Lab README and paper, plus the AI4Chem README and CheMatAgent paper. DrugAgent’s architecture and deployment limitations come from paper v2. SciAgents’ licence discrepancy is visible in its licence file and package metadata.