Overview
PubChem MCP (cyanheads), packaged as @cyanheads/pubchem-mcp-server, is a community MCP integration for retrieving chemical information through PubChem's PUG REST and PUG View APIs. It is an interface to the upstream database, not PubChem itself or a predictive chemistry model. Its 10 tools and six URI-templated resources support compound identification, record enrichment and biological-target exploration within MCP-enabled workflows.
Inputs include compound names, SMILES, InChIKey, molecular formulas and structure queries; most record-retrieval tools use PubChem compound identifiers (CIDs). Assay searches accept gene symbols, protein names, Gene IDs or UniProt accessions. Outputs include structured properties, descriptions, synonyms, GHS classifications, assay activity records, interactions and external identifiers. Structure retrieval provides base64-encoded PNG diagrams or 3D conformers as parsed atom-and-bond JSON or raw V2000 SDF. Optional Lipinski and Veber assessments are calculated from retrieved properties rather than predicted by a trained model.
The documented deployment choices are a local STDIO process, a self-hosted Streamable HTTP server and a public hosted endpoint. Calls are read-only, with request queuing, retries and cancellation handling. Structured statuses distinguish missing records from absent data, while pagination and truncation fields help clients handle bounded responses.
Coverage remains dependent on PubChem records: some compounds lack GHS information or computed 3D coordinates. Search and enrichment operations have explicit batch and pagination limits; descriptions and classification are fetched for only the first 10 CIDs in a batch. Some MCP clients expose tools but not resources, so resource access should not be assumed when planning an integration.
Key Features
- Search compounds by name, SMILES, InChIKey, Hill formula, substructure, superstructure or 2D Tanimoto similarity, with optional property enrichment and unresolved-identifier reporting.
- Retrieve selected physicochemical properties for up to 100 CIDs, with optional synonyms, descriptions, pharmacological classification and property-derived Lipinski/Veber assessments.
- Fetch 2D PNG structure diagrams and 3D conformers as atom-and-bond JSON or V2000 SDF, with explicit preview caps and a distinct error for unavailable 3D structures.
- Retrieve attributed GHS safety records for batches of up to 25 CIDs, distinguishing missing classifications from nonexistent compound records.
- Explore bioactivity by outcome or molecular target, search assays by target identifiers, and retrieve entity summaries, interactions and external database cross-references.
- Expose read-only tools and URI-templated resources over STDIO or Streamable HTTP, with queued requests, retry handling, cancellation propagation and disclosed partial failures.
Use Cases
- Suggested evaluation: resolve a mixed list of compound names and SMILES to CIDs, inspect unresolved inputs, and enrich matched records with selected physicochemical properties.
- Suggested evaluation: gather GHS classifications and literature or patent cross-references for a compound inventory, retaining source attribution and missing-data statuses.
- Suggested evaluation: explore a target's PubChem bioassays and filter a compound's recorded activity profile by outcome or target before reviewing individual assay summaries.
- Suggested evaluation: retrieve available 3D conformers for a structure-processing workflow, checking missing-coordinate errors and preview truncation before downstream use.
How to Use
- Read the project README to choose between hosted Streamable HTTP, local STDIO and self-hosted HTTP. Check the documented runtime prerequisites if running locally.
- For hosted access, configure your MCP client with the documented Streamable HTTP endpoint, https://pubchem.caseyjhand.com/mcp. For local deployment, follow the README's package or container configuration rather than substituting undocumented options.
- Begin an evaluation with
pubchem_search_compoundsand a small set of known names, SMILES or InChIKey inputs. Supply the fields required by the selected strategy and inspectunresolvedIdentifiersand match counts. - Pass resolved CIDs to
pubchem_get_compound_details, selecting properties needed for your workflow. Add safety, cross-reference or structure tools as appropriate; distinguish absent data from missing records. - For biological exploration, use
pubchem_search_assayswith a supported target identifier and inspect relevant summaries or bioactivity pages. Follow offsets and truncation indicators, then compare representative outputs with the upstream records before adopting the integration.