Separate the data source from the integration
A chemical-data agent needs more than a searchable database: it needs a reliable boundary between the source record, the interface that retrieves it and the explanation it produces. This collection compares four candidates at those different layers. It is a planning aid, not a compatibility matrix or evidence that the resources have been executed together.
For every proposed connection, record the upstream service, repository maintainer, selected source revision and deployment choice. The cyanheads servers are community integrations; their names do not establish ownership by PubChem or ChEMBL, or a provider service guarantee. Keep software licensing separate from underlying database access and attribution terms.
Start with a decision question: Does the task require compound identification, measured bioactivity, materials data access or procedural guidance? Choose the smallest interface set that answers it rather than connecting everything by default.
Match each resource to its role
- PubChem MCP (cyanheads) wraps PubChem's PUG REST and PUG View APIs. Its documented scope includes identifier and structure searches, physicochemical properties, attributed GHS records, bioactivity, cross-references and 3D conformers. Most record retrieval uses compound identifiers, or CIDs. Property-derived Lipinski and Veber assessments are calculations from retrieved fields, not predictions from a trained model.
- ChEMBL MCP (cyanheads) connects molecule and target searches to measured activity records, assay provenance, drug mechanisms and indications. Bioactivity queries accept a molecule, a target or a compound–target pair. Optional DuckDB-backed DataCanvas functionality supports analysis of larger activity sets.
- Materials Project API client is the client project distributed as
mp-api. The prepared resource reference records an optional MCP dependency group and thempmcpentry point. That establishes an integration entry point, not its exposed tools or a verified setup. Query fields, authentication and response contracts remain unspecified in the source excerpts and need confirmation before pipeline design. - Scientific Agent Skills supplies procedural instructions and supporting assets, rather than an independent database or standalone research agent. Its Database Lookup guidance covers endpoint selection, pagination, access requirements and provenance, including PubChem and ChEMBL. Individual skills may include helpers or templates; dependencies and credentials are arranged separately.
The MCP servers document STDIO and Streamable HTTP deployment choices. Shared transport support alone does not prove that a particular host exposes every tool or resource correctly.
Define the input and output contract
Before allowing autonomous calls, write a short contract for each task:
- Inputs: permitted identifier types, structure representations, organism or target restrictions, requested properties and result bounds.
- Outputs: source identifiers, selected fields, units where applicable, query filters, retrieval time, completeness indicators and source attribution.
- Continuation: how pagination is followed, when retrieval stops and how an incomplete set is labelled.
- Recovery: which failures allow a corrected query and which require human clarification.
For PubChem, distinguish unresolved inputs, nonexistent CIDs and existing compounds without a requested measurement. For ChEMBL, preserve missing numeric values as null rather than zero. Keep the ranked and null_potency views distinct.
Ask: Can the agent tell whether it received a complete dataset, a preview or a filtered empty result? ChEMBL's inline preview and staged tables have separate bounds. PubChem exposes pagination and truncation information, and some enrichment fields have batch limits. Require the agent to retain those disclosures in its final answer.
Worked planning example: a hypothetical compound–target lookup
Suppose a researcher wants an evidence table for three public compound names and one protein target. This is a hypothetical plan, not a reported test or scientific result.
Proposed input: three names, a researcher-supplied UniProt accession, an organism restriction and one activity measurement type, such as IC50.
Proposed sequence:
- Use
pubchem_search_compoundsto resolve the names. Stop for clarification if several candidate records remain; do not silently select the first result. - Retrieve selected properties and retain CIDs and structure identifiers as identity evidence.
- Independently resolve ChEMBL molecules using a supported identifier or structure input. Treat the PubChem-to-ChEMBL mapping as a proposed integration boundary requiring identity checks, not an automatic equivalence.
- Use
chembl_search_targetswith the accession and organism restriction, then querychembl_get_bioactivitiesfor each resolved compound–target pair. - Inspect relevant assays with
chembl_get_assaybefore proposing comparisons.
Planned output: one table linking original inputs to resolved records, and another retaining activity type, value, units, assay and target identifiers, filters and completeness status. Unresolved mappings and absent measurements should appear explicitly.
Decision questions include: Are the structures consistent across sources? Is the target the intended organism and target type? Are the assays comparable? The ChEMBL documentation limits potency comparisons to one standard_type; matching that field is not, by itself, sufficient scientific justification for ranking every assay together.
Evaluate permissions and failure handling
The following are suggested evaluation steps, not completed checks.
Inventory network destinations, credentials, logging, storage and local writes for the chosen configuration. The PubChem and ChEMBL integrations describe keyless upstream access, but deployment authentication and host permissions are separate questions. Read-only retrieval also does not mean that logging or dataset staging leaves no local artifacts.
Inspect only the Scientific Agent Skills subset needed for the task. Skills can direct package installation, code execution, network requests and file changes. Check their instructions and helpers before granting those permissions.
Use non-sensitive inputs to exercise ambiguous names, malformed structures, missing records, provider failures and pagination. For ChEMBL canvas use, also assess staging limits and disabled-provider behavior. Save the exact configuration and source revision if testing later occurs; leave unsupported pairings unknown.
Preserve provenance and acknowledge limitations
Compare the agent's final prose with the retrieved records. A fluent summary must not turn missing GHS data into an absence-of-hazard claim, null potency into inactivity, or a truncated preview into a complete search.
Coverage depends on upstream records. PubChem may lack classifications or computed 3D coordinates; ChEMBL activity interpretation depends on assay context. The supplied Materials Project evidence does not establish property coverage or authentication behavior. Skill instructions do not establish scientific validity or live-service reliability.
This collection remains an unpublished draft pending source review and editorial selection.
Short actionable checklist
- Choose a task and the minimum resource set.
- Record upstream provider, maintainer, revision and deployment.
- Specify identifiers, fields, units and stopping rules.
- Preserve missing-data, error and truncation distinctions.
- Inspect permissions, logs, storage and skill helpers.
- Evaluate identity mappings and assay comparability with public examples.
- Retain source records alongside the agent summary; document only checks actually performed.