Skip to content

ChemAgent (Gerstein Lab)

ChemAgent is a chemical-reasoning research framework that decomposes problems and retrieves reusable memories, with released code for constructing memory pools and running SciBench experiments.

Catalog updated ·

Overview

ChemAgent (Gerstein Lab) is a research framework for investigating how a self-updating library can support large language models on multi-step chemical reasoning problems. The paper describes a workflow that breaks chemical tasks into smaller tasks and stores those experiences in a structured library for subsequent queries. Rather than supplying a standalone chemistry prediction model, ChemAgent adds memory-supported decomposition and solution generation around an LLM.

The framework takes chemical reasoning problems as inputs, retrieves relevant stored information, and refines that information while developing solutions. Its documented architecture includes task splitting, execution, association, and reflection modules. The paper describes three memory types; the repository instructions identify Plan Memory, Execution Memory, and Knowledge Memory. Configuration settings select memory storage locations and control strategy refinement after execution failure and whether evaluation is enabled.

The repository README documents code for building memory pools and running experiments on the SciBench-derived datasets atkins, chemmc, matter, and quan. It also identifies directories containing experimental outputs and stored memories. Memory construction is a prerequisite in the development workflow, not an automatic installation step.

This is released research code rather than a documented turnkey service. The installation section remains a placeholder, and the supplied excerpts do not establish a complete dependency setup or current model compatibility. Drug discovery and materials science appear in the paper as potential future applications, not demonstrated workflows.

Key Features

  • Decomposes chemical reasoning tasks into subtasks that can populate a reusable, self-updating library.
  • Retrieves and refines relevant memories to support decomposition and solution generation for new problems.
  • Separates task splitting, execution, association, and reflection into functional modules.
  • Provides configurable Plan Memory and Execution Memory storage locations, with a documented Knowledge Memory activation setting.
  • Documents development mode for memory construction and test mode for experiments on `atkins`, `chemmc`, `matter`, and `quan`.
  • Exposes configuration controls for strategy refinement after execution failure and for enabling evaluation.

Use Cases

  • Intended evaluation: reproduce the documented SciBench chemical-reasoning workflow after resolving the incomplete environment setup.
  • Intended evaluation: investigate how reusable subtask memories affect solution generation on the supplied reasoning datasets.
  • Intended evaluation: compare runs with strategy refinement or evaluation settings changed, checking generated solutions against dataset references.

How to Use

  1. Read the ChemAgent paper to understand task decomposition, memory construction, and the experimental scope before attempting reproduction.
  2. Inspect the released repository and its README. Identify the datasets, memory directories, and experimental outputs; resolve dependencies separately because installation instructions are unfinished.
  3. Review assets/config.yml for model settings, Plan Memory and Execution Memory paths, refinement, and evaluation. Verify the Knowledge Memory setting in the code: the README shows image in its example but describes img in the explanation.
  4. Follow the README’s development workflow: select two examples from fewshot_examples in XAgent/agent/simple_agent/prompt.py, adjust configuration, and use the documented development mode to construct memory for the selected dataset.
  5. Use the documented test mode in dev_test.py with one of atkins, chemmc, matter, or quan. As an evaluation step, inspect generated solutions and experimental outputs for calculation and reasoning errors rather than assuming successful scientific validation.

Related resources

Allegro

Model

Allegro implements an E(3)-equivariant interatomic potential as a NequIP extension, with documented GPU acceleration options and a separate plugin for LAMMPS simulations.

Open sourcePython

Computational Chemistry · Materials Discovery

ASE

Open Source

ASE is a Python atomistic simulation library connecting conventional codes and external machine learning potentials through a common calculator interface.

Open sourcePython

Computational Chemistry

An ASE routing Skill in the computational-chemistry-agent-skills collection that separates workflow preparation from calculator configuration and delegates execution elsewhere.

Computational Chemistry · Materials Discovery

Cantera Skill (K-Dense) provides instructions and a Python helper for homogeneous ignition calculations, with mechanism provenance, conservation diagnostics, and numerical refinement checks.

Open sourcePython

Computational Chemistry · Process Optimization

ChemAgent (AI4Chem) is a research framework for chemistry and materials tool use, linked to the CheMatAgent paper on tree-search planning, tool execution and ChemToolBench-based training.

Computational Chemistry · Materials Discovery

ChemGraph is a Python agent framework that connects natural-language chemistry requests to molecular construction, simulations, analysis, and reporting, with CLI, Python, Streamlit, and MCP interfaces.

Open sourcePython

Computational Chemistry · Materials Discovery

Related guides

Agents

AI Agents for Chemistry: From Chatbots to Autonomous Research

A practical guide to eight chemistry and biomedical research agents: how tools, memory, planning and multi-agent roles work, which projects fit different tasks, and how to evaluate bounded autonomy with scientific oversight.

Workflows

AI Literature Research for Chemistry

Build a traceable chemistry literature workflow with PubMed MCP: refine searches, inspect metadata and abstracts, synthesize multiple papers, follow citations, and separate retrieved evidence from agent-generated hypotheses.