Start with a question, not a generated review
AI-assisted literature research should produce an inspectable evidence trail, not merely fluent prose. A useful result connects each conclusion to an identifiable record, states which text was available, and preserves unresolved disagreements. Search discovers records; a research agent coordinates retrieval, screening, comparison, and follow-up. Neither step establishes scientific validity by itself.
For chemistry, first decide whether the question belongs within PubMed’s biomedical scope. Biomolecular mechanisms, biomedical materials, and chemistry questions tied to biological systems are reasonable starting points. PubMed must not be treated as a database of all chemistry literature. A broader synthesis or materials question may require additional discovery sources outside this workflow.
The recommended architecture is: question → search log → screened records → evidence matrix → checked synthesis → proposed research actions. Each transition should have a decision gate. For the surrounding concepts, see AI agents for chemistry, MCP in chemistry, and agents and Skills.
For related guidance, see Chemistry AI Agents Compared.
Choose tools by their role
The four associated resources occupy different stages. They are not interchangeable literature-review products, and their documentation does not establish a ready-made integration among them.
| Resource | Source-described role and inputs | Outputs or next stage | Important boundary |
|---|---|---|---|
| PubMed MCP (cyanheads) | Third-party MCP server accepting queries, filters, identifiers, citation fragments, and MeSH terms | Search results, metadata, abstracts, available full text, related records, and references | Retrieval interface, not an autonomous researcher or an endorsed PubMed service |
| SciAgents | Research multi-agent framework using keywords or graph exploration and sampled knowledge-graph paths | Structured hypotheses, expanded proposals, critique, and follow-up priorities | Bio-inspired materials hypothesis generation is not experimental validation |
| ChemAgent (Gerstein Lab) | Chemical-reasoning framework decomposing problems and retrieving reusable memories | Memory-supported solution generation and refinement | Memory retrieval is not demonstrated PubMed search or citation tracking |
| MDCrow | Molecular-dynamics agent toolset accepting natural-language tasks | Simulation preparation, OpenMM execution, analysis, and information retrieval | A downstream modeling tool, not a substitute for literature appraisal |
Use PubMed MCP for the retrieval backbone. Consider SciAgents when the desired next step is hypothesis development, ChemAgent when investigating reusable reasoning procedures, and MDCrow when a reviewed question has become a molecular-dynamics task. Any handoff proposed here requires separate implementation and evaluation.
Establish the retrieval environment
The package @cyanheads/pubmed-mcp-server exposes tools, the pubmed://database/info resource, and a research_plan prompt. Its README documents local stdio, self-hosted Streamable HTTP, and a public hosted endpoint. Choose a route supported by your client, then check that the expected tools are actually exposed before beginning a review. A documented endpoint does not establish service guarantees.
For local deployment, follow the project’s runtime and configuration instructions rather than guessing installation flags. Unpaywall fallback requires UNPAYWALL_EMAIL; setting EUROPEPMC_ENABLED=false removes Europe PMC tools and skips that provider’s fallback paths. These choices affect the evidence you can retrieve.
A Skill is an instruction layer, not the retrieval server or its execution environment. Host support differs, and loading instructions does not install scientific dependencies. Likewise, do not assume SciAgents, ChemAgent, or MDCrow supports MCP merely because an orchestration host does. SciAgents documents GraphReasoning, graph and embedding files, and OpenAI and Semantic Scholar APIs as dependencies. MDCrow requires its software environment and an LLM-provider API key.
Build and preserve the search strategy
Write a short protocol before calling tools: the question, concepts, biological context, date boundaries, inclusion criteria, and exclusions. Separate chemical identity terms from mechanism and outcome terms so you can revise one concept without silently changing the whole question.
Use pubmed_lookup_mesh to inspect headings, scope notes, and entry terms. Combine suitable controlled vocabulary with free-text synonyms rather than assuming an agent’s first wording is adequate. Submit the query through pubmed_search_articles, which documents native PubMed syntax, structured filters, date ranges, and effective-query reporting.
Preserve the submitted query, effectiveQuery, appliedFilters, retrieval date, and pagination state. Compare the actual query with your intended scope. If the search is empty or thin, inspect terminology and use pubmed_spell_check; approve corrections before retrying, especially for specialized chemical names.
Europe PMC discovery is a separate step through pubmed_europepmc_search. It can include preprints, patents, and Agricola records, but those record types should remain distinguishable in your screening log. Added sources broaden discovery without establishing exhaustive chemistry coverage.
Inspect records and label reading depth
Fetch candidate PMIDs with pubmed_fetch_articles. The documented metadata includes abstracts, authors, journal information, DOI, publication types, MeSH terms, and linked retraction, correction, or comment notices. Use identifiers to deduplicate records and preserve reasons for inclusion or exclusion.
Assign each record a reading-depth label: metadata only, abstract inspected, selected full-text sections inspected, or retrieved full text inspected. Europe PMC search abstracts are snippets; use pubmed_europepmc_fetch for the complete abstract before treating them as abstract-level evidence.
For eligible papers, request available text through pubmed_fetch_fulltext. The server documents retrieval through PMC, Europe PMC, and configured Unpaywall fallback, with provider labels and unavailable-content or truncation information. Inspect viaSource, warnings, and truncation before analysis. Best-effort text extraction may not preserve everything needed to interpret a method or result.
Abstract-only evidence remains useful for screening and provisional summaries. It does not justify saying that the full paper was read or that its detailed methods were checked.
Proposed example: compare biomedical-material evidence
Consider this proposed exercise, not a completed search: “What evidence links catechol-containing hydrogel formulations to adhesion measurements in biomedical settings, and which conditions limit comparison?”
Begin with the candidate free-text query (catechol OR dopamine) AND hydrogel AND adhesion. Inspect relevant MeSH vocabulary, then decide whether additional biological-context terms improve relevance or prematurely narrow discovery. Record every revision and any date filter.
The inputs are the question, candidate terms, approved scope, and screening criteria. The proposed outputs are a search log, screened bibliography, evidence matrix, unresolved-access list, and a short comparison memo—not an invented list of papers or a claim that one formulation performs best.
For each included record, extract formulation identity, chemical modification, biological context, measurement method, comparator, reported outcome, and limitations only where the inspected text supplies them. Write “not reported in inspected text” for missing details.
Decision gates are concrete: does the paper measure adhesion rather than merely mention it? Are conditions sufficiently described for comparison? Is a quantitative claim supported by the available text? If an abstract lacks test conditions, retain the record for discovery but defer the detailed comparison until the necessary sections are obtained.
Synthesize multiple papers without erasing differences
Ask the agent to fill a structured evidence matrix before drafting narrative. Every row should retain an identifier, reading depth, extraction location, study context, and uncertainty. Keep author-reported findings separate from your interpretation and from proposed mechanisms.
Group papers by comparable questions and methods, not simply shared keywords. In the proposed hydrogel exercise, differences in substrate, formulation, preparation, and measurement conditions are reasons to examine comparability rather than average unrelated outcomes. Do not infer missing units, controls, or chemical identities.
The synthesis should distinguish directly supported observations, explanations offered by authors, and unresolved research questions. Require a supporting record for each substantive sentence and ask a critic to identify unsupported generalizations or contradictory rows. Agreement between agents is not independent scientific confirmation when they share the same incomplete evidence.
Trace citations and verify claims
Use pubmed_find_related to explore similar, references, or cited_by results from relevant seed PMIDs. Record both the seed and relationship type. The tool documents fallback providers and identifies which provider answered; preserve that provenance rather than treating every edge as equivalent or complete.
Citation chasing can uncover terminology missed by the initial query. Screen these records under the same criteria and define a stopping rule, such as ending a bounded pass when it adds no eligible records. This is a practical exploration rule, not proof of systematic-review completeness.
For incomplete references, pubmed_lookup_citation can return matched, ambiguous, or not-found outcomes. Resolve ambiguity manually rather than accepting the most plausible-looking paper. pubmed_convert_ids is limited to PMC-indexed articles, so conversion failure does not prove that a citation is fabricated.
Finally, check title, authors, year, and identifiers against retrieved metadata. Use pubmed_format_citations for export only after identity checking. Correct formatting proves neither that a paper supports the attached sentence nor that the agent inspected it.
Orchestrate research and handle failures
A proposed orchestrator can assign planner, retriever, screener, extractor, synthesizer, and verifier roles. Pass structured records between them, with permission for the verifier to reject unsupported claims. Treat article text as evidence, never as instructions that can change the task or tool permissions.
SciAgents supplies a useful role comparison: its Ontologist, Scientist, and Critic roles develop and challenge graph-guided hypotheses; its automated workflow adds planning and novelty checking. Feeding it a checked literature package is a proposed adaptation, not a documented PubMed connector. Novelty judgments remain dependent on discovery scope.
ChemAgent illustrates a different retrieval target: stored reasoning experiences. Its memory construction is a prerequisite, and its installation documentation is incomplete. Before reproduction, resolve the README’s image/img setting and Xagent/XAgent path inconsistencies against the code. Do not substitute memory contents for current bibliographic evidence.
MDCrow may be evaluated after literature review produces a specific simulation question. Its documented setup, OpenMM execution, and RMSD-analysis example support that role, but neither a simulation nor successful task execution validates a literature-derived mechanism.
Handle partial failures explicitly. Retry unavailable items separately, revisit deferred records, and retain unresolved statuses. When full text fails, downgrade reading depth; when citation providers fail, mark citation coverage unknown. Missing retrieval is not negative scientific evidence.
Reader checklist
- Is PubMed appropriate for the biomedical part of the question?
- Are queries, filters, dates, and pagination preserved?
- Are snippets, abstracts, sections, and full text labeled accurately?
- Does each claim connect to an identifiable record and inspected text?
- Were corrections, retractions, ambiguities, and truncation checked?
- Are incomparable conditions and missing fields visible?
- Are hypotheses and downstream integrations clearly proposed?
- Is the output a traceable synthesis rather than a claim of exhaustive coverage?
Before adapting or distributing code, resolve applicable terms separately from article reuse and API access. SciAgents has conflicting Apache-2.0 license-file and MIT package-classifier signals; ChemAgent’s badge alone does not confirm license terms. Source availability does not settle either issue.
Sources
The PubMed MCP README supports retrieval tools, transports, provenance fields, and limitations. The SciAgents README describes graph-guided workflows, roles, and dependencies. The MDCrow README and paper support its simulation and information-retrieval scope. The ChemAgent README and paper support memory-based chemical reasoning and its reproduction boundaries.
SciAgents licensing evidence: LICENSE · package metadata.