What changes when a chatbot becomes an agent?

A useful distinction is whether the system merely proposes an answer or can act on a task, inspect the result and decide what to do next. Explaining a molecular dynamics protocol is conversational assistance. Selecting preparation tools, launching a simulation and requesting an analysis is an agent workflow.

Here, autonomous research means a spectrum of delegated capabilities, not a promise of unattended scientific discovery. A researcher may delegate retrieval, planning, calculation or implementation while retaining authority over assumptions, execution permissions and conclusions. More autonomy should mean a clearer task contract—not fewer scientific checks.

This guide focuses on eight chemistry and biomedical research projects. For broader distinctions among models and infrastructure, use the chemistry AI stack guide and the MCP, Skills, agents and APIs guide. The separate agents and Skills guide addresses cross-science workflows and permission assessment.

The working parts: model, tools, memory, planning and execution

Read an agent architecture as a division of responsibilities:

  • LLM: interprets objectives, proposes steps and explains results. Its text is not itself a calculation or measurement.
  • Tools: retrieve records, construct structures, predict properties, run simulations or manipulate files. Their scientific scope comes from the underlying method.
  • Memory: preserves reusable information, such as earlier task solutions or workflow templates. It needs provenance and inspection before reuse.
  • Planning: turns an objective into tool choices, parameters and dependencies; some frameworks explore alternative plans.
  • Execution: performs the actions and returns artifacts, errors or reports that inform the next decision.

These parts are not interchangeable products. An orchestration framework is not a standalone property model. A wrapper or tool interface exposes another program’s capabilities; it does not replace that program. A hosted model API supplies model access, not the local scientific environment.

The official Agent Skills format packages instructions in SKILL.md, optionally with supporting resources or scripts. Host support differs, and loading instructions does not install execution dependencies. ChemGraph’s Skills and STELLA’s workflow templates should therefore be understood as supporting procedures, not additional autonomous researchers.

MCP defines context exchange between hosts, clients and servers. ChemGraph documents both serving and consuming MCP tools. The current protocol overview does not establish compatibility between every older project implementation and a current host.

For the tool-access layer, see MCP for chemistry.

Eight projects compared by research role

The table describes source-supported roles, not a ranking or tested interoperability matrix. The entry points are suggestions for evaluation.

Project Source-described role and outputs Suggested entry point and main boundary
ChemGraph Orchestrates molecular construction, ASE calculations, analysis and reporting; artifacts can include structures, trajectories, spectra and results Start with single_agent; usable calculators depend on the environment, and EMT is for setup checks rather than general high-accuracy chemistry
MDCrow Coordinates molecular dynamics preparation, OpenMM execution, analysis and scientific information retrieval Inspect a small simulation-and-RMSD workflow; the supplied evidence does not establish comprehensive settings or output formats
ChatMOF Routes MOF requests to retrieval, property prediction and structure generation tools Begin with search; local prediction needs module setup, and generation additionally requires GRIDAY
SciAgents Uses graph context and specialized agents to generate structured hypotheses and expanded research proposals Examine a sampled graph path and its resulting proposal; outputs remain hypotheses requiring follow-up
ChemAgent (Gerstein Lab) Decomposes chemical reasoning problems and retrieves reusable Plan, Execution and Knowledge memories Reproduce a documented SciBench-derived task after constructing memory; installation guidance is incomplete
ChemAgent (AI4Chem) Framework linked to the CheMatAgent paper on chemistry/materials tool learning and tree-search planning Study tool selection and parameter filling; deployment requirements and checkpoint availability are not established
STELLA Coordinates biomedical literature, databases and analyses using manager, development and critic roles, reusable templates and optional tool creation Inspect an evidence-backed ranking or analysis task; local execution requires OpenRouter access and task-dependent resources
DrugAgent (Liu et al.) Paper-described Planner–Instructor collaboration for implementing and evaluating drug-discovery ML workflows Use the methodology as a reproduction starting point; supplied repository evidence does not establish runnable setup or a code licence

Two similarly named entries require care: the Gerstein Lab ChemAgent studies memory-supported chemical reasoning, whereas the AI4Chem repository identifies a ChemAgent framework associated with the CheMatAgent paper. Neither name establishes equivalent functionality.

Match the agent to the scientific task

For computational chemistry, separate orchestration from the engine doing the calculation. ChemGraph offers a broader construction–calculation–reporting layer, while MDCrow is specifically oriented toward molecular dynamics. Choose according to the calculation and artifacts needed, then check whether the required engine and settings are available.

For materials discovery, ChatMOF addresses MOF retrieval, prediction and generation. SciAgents instead develops bio-inspired materials hypotheses from graph relationships. A generated structure and a research proposal answer different questions: one supplies a computational candidate, the other supplies a direction to investigate.

For molecular design, define whether the deliverable is a constructed structure, a predicted property or a candidate aimed at a target property. ChatMOF explicitly describes property-directed MOF generation. ChemGraph’s molecular construction should not be silently upgraded into a general inverse-design capability.

For literature research, STELLA offers biomedical search and evidence-oriented workflows; SciAgents uses retrieved information within hypothesis development. ChemGraph also documents PDF/text retrieval workflows, and MDCrow’s paper describes literature and database retrieval. Require traceable supporting sources rather than accepting a fluent summary alone.

For drug discovery, DrugAgent’s role is ML programming, not laboratory discovery. Its paper studies PAMPA, HIV and DAVIS classification tasks. STELLA describes biomedical prioritization and virtual-screening tools, while ChemGraph has a dependency-specific docking workflow. These stages are not substitutes for one another or for experimental validation.

When multiple agents help—and what their roles mean

Multiple agents are useful when responsibilities can be separated and their handoffs inspected. SciAgents assigns concept clarification, proposal development, refinement and criticism to different roles; its automated workflow also adds planning and novelty checking. STELLA separates management, development and critique, with optional tool creation. DrugAgent divides strategy exploration from domain-guided implementation.

The important question is not how many agents are present, but what evidence crosses each boundary. A planner should hand over explicit inputs and constraints. An executor should return artifacts and failure information. A critic should identify unsupported assumptions, not merely rewrite the answer more confidently.

For evaluation, prefer a single-agent starting point when one tool sequence is sufficient. Add specialist roles only when they provide a checkable benefit, such as independent inspection of parameters or comparison of alternative implementations. Agent agreement should not be treated as independent scientific confirmation.

Proposed example: a bounded molecular dynamics evaluation

This is a proposed evaluation, not a completed test. Use MDCrow’s documented protein-simulation-and-RMSD example as the starting task, but ask for an inspectable workflow rather than only a final answer.

Inputs: a researcher-approved protein identifier or structure, temperature, a deliberately short run duration, the requested RMSD analysis, an isolated software environment and a compute limit. Have the researcher specify preparation and simulation choices that are not established by the example.

  1. Plan: request a list of preparation, execution and analysis steps. Decision point: are required settings explicit, or has the agent supplied unreviewed defaults?
  2. Approve: inspect the structure and proposed setup before execution. If a parameter is ambiguous or unsupported, stop and resolve it rather than permitting substitution.
  3. Execute: authorize only the bounded run. Preserve generated inputs, execution logs and available simulation artifacts; confirm their actual formats locally.
  4. Analyze: request RMSD over time with the atom selection and reference stated. Decision point: do the analysis choices match the scientific question?
  5. Audit: compare the summary with the saved evidence. Classify missing artifacts, execution errors and unsupported conclusions separately.

Requested outputs: a parameter record, execution status, artifact inventory, RMSD analysis and a short account of unresolved assumptions. This is an evaluation contract, not a claim that MDCrow automatically produces every item.

A successful short run would establish workflow completion under those conditions—not adequate sampling or a biological conclusion. Any downstream connection to DrugAgent for ML workflow development would be a proposed integration, requiring separately defined data, labels, evaluation rules and handoffs; no such interoperability is established here.

Failure handling and scientific limits

Use failures to narrow the task, not to encourage increasingly unrestricted retries.

  • Missing engine or dependency: stop the affected stage and identify what is absent. Do not substitute a calculator with a different scientific role without approval.
  • Malformed parameters or structures: inspect the tool call and input artifact before retrying; preserve the failed attempt.
  • Missing citations or contradictory evidence: return an unresolved evidence item, not an invented reconciliation.
  • Failed calculation or analysis: distinguish execution failure from an unfavorable scientific result. A polished report cannot repair missing computation.
  • Unsafe file or code action: require review and isolation. ChemGraph explicitly states that workspace shell access is not confined to the workspace.

Scientific review should examine identities, units, calculation settings, convergence, analysis definitions and whether the method supports the conclusion. For proposed ML evaluations, also inspect data provenance, splits and label handling; the reproducible chemistry AI guide provides a broader framework.

Memory and templates introduce another decision: should this result become reusable guidance? Recommend admitting only inspected records with provenance and clear scope. STELLA notes that rankings can vary with model stochasticity and literature-index changes. SciAgents’ novelty assessments remain search-dependent judgments. DrugAgent’s authors explicitly warn against direct deployment in a drug-discovery pipeline because fabricated results could misdirect follow-up work.

Selection checklist and unresolved setup questions

Before choosing a project, check:

  • Is the intended output an answer, hypothesis, structure, simulation artifact or implemented ML solution?
  • Are required engines, models, resources and credentials documented?
  • Can you inspect tool parameters and outputs separately from narrative summaries?
  • Are execution permissions and stopping conditions explicit?
  • Can failed attempts be preserved without overwriting useful evidence?
  • Who approves scientific assumptions and downstream use?
  • Are code, weights, data and service terms assessed separately?

Documentation gaps affect selection directly. ChatMOF’s online demonstration is search-focused; it is not a general prediction-and-generation trial. Gerstein Lab ChemAgent’s image/img setting and Xagent/XAgent path inconsistencies should be resolved against the code before reproduction. AI4Chem supplies no deployment instructions in the provided README. STELLA’s current README and linked arXiv paper represent different presentations, so reproduction should identify the snapshot being followed.

SciAgents also has conflicting licence metadata: its licence file contains Apache License 2.0 while package metadata declares MIT. Resolve that before reuse or redistribution. Repository visibility alone does not establish licensing; consult the software, weights and data licences guide for the separate layers.

Continue with these related guides: Chemistry AI Agents Compared: Capabilities, Code Availability and Research Scope.

Sources

Primary capability sources are the ChemGraph README, MDCrow README and paper, ChatMOF README, and SciAgents README.

Memory and tool-learning claims follow the Gerstein Lab ChemAgent README and paper, and the AI4Chem README and CheMatAgent paper. Biomedical orchestration follows the STELLA README; DrugAgent architecture and deployment cautions follow its paper.

Infrastructure distinctions follow the official MCP architecture overview and Agent Skills overview. The SciAgents licensing conflict is visible in its licence file and package metadata.