Overview
ChemAgent (Gerstein Lab) is a research framework for investigating how a self-updating library can support large language models on multi-step chemical reasoning problems. The paper describes a workflow that breaks chemical tasks into smaller tasks and stores those experiences in a structured library for subsequent queries. Rather than supplying a standalone chemistry prediction model, ChemAgent adds memory-supported decomposition and solution generation around an LLM.
The framework takes chemical reasoning problems as inputs, retrieves relevant stored information, and refines that information while developing solutions. Its documented architecture includes task splitting, execution, association, and reflection modules. The paper describes three memory types; the repository instructions identify Plan Memory, Execution Memory, and Knowledge Memory. Configuration settings select memory storage locations and control strategy refinement after execution failure and whether evaluation is enabled.
The repository README documents code for building memory pools and running experiments on the SciBench-derived datasets atkins, chemmc, matter, and quan. It also identifies directories containing experimental outputs and stored memories. Memory construction is a prerequisite in the development workflow, not an automatic installation step.
This is released research code rather than a documented turnkey service. The installation section remains a placeholder, and the supplied excerpts do not establish a complete dependency setup or current model compatibility. Drug discovery and materials science appear in the paper as potential future applications, not demonstrated workflows.
Key Features
- Decomposes chemical reasoning tasks into subtasks that can populate a reusable, self-updating library.
- Retrieves and refines relevant memories to support decomposition and solution generation for new problems.
- Separates task splitting, execution, association, and reflection into functional modules.
- Provides configurable Plan Memory and Execution Memory storage locations, with a documented Knowledge Memory activation setting.
- Documents development mode for memory construction and test mode for experiments on `atkins`, `chemmc`, `matter`, and `quan`.
- Exposes configuration controls for strategy refinement after execution failure and for enabling evaluation.
Use Cases
- Intended evaluation: reproduce the documented SciBench chemical-reasoning workflow after resolving the incomplete environment setup.
- Intended evaluation: investigate how reusable subtask memories affect solution generation on the supplied reasoning datasets.
- Intended evaluation: compare runs with strategy refinement or evaluation settings changed, checking generated solutions against dataset references.
How to Use
- Read the ChemAgent paper to understand task decomposition, memory construction, and the experimental scope before attempting reproduction.
- Inspect the released repository and its README. Identify the datasets, memory directories, and experimental outputs; resolve dependencies separately because installation instructions are unfinished.
- Review
assets/config.ymlfor model settings, Plan Memory and Execution Memory paths, refinement, and evaluation. Verify the Knowledge Memory setting in the code: the README showsimagein its example but describesimgin the explanation. - Follow the README’s development workflow: select two examples from
fewshot_examplesinXAgent/agent/simple_agent/prompt.py, adjust configuration, and use the documented development mode to construct memory for the selected dataset. - Use the documented test mode in
dev_test.pywith one ofatkins,chemmc,matter, orquan. As an evaluation step, inspect generated solutions and experimental outputs for calculation and reasoning errors rather than assuming successful scientific validation.