Skip to content

PubChem MCP (cyanheads)

Casey Hand

PubChem MCP (cyanheads) connects MCP clients to PubChem compound and bioassay data, with identifier and structure searches, property retrieval, safety records, cross-references and 3D conformers.

Catalog updated ·

Overview

PubChem MCP (cyanheads), packaged as @cyanheads/pubchem-mcp-server, is a community MCP integration for retrieving chemical information through PubChem's PUG REST and PUG View APIs. It is an interface to the upstream database, not PubChem itself or a predictive chemistry model. Its 10 tools and six URI-templated resources support compound identification, record enrichment and biological-target exploration within MCP-enabled workflows.

Inputs include compound names, SMILES, InChIKey, molecular formulas and structure queries; most record-retrieval tools use PubChem compound identifiers (CIDs). Assay searches accept gene symbols, protein names, Gene IDs or UniProt accessions. Outputs include structured properties, descriptions, synonyms, GHS classifications, assay activity records, interactions and external identifiers. Structure retrieval provides base64-encoded PNG diagrams or 3D conformers as parsed atom-and-bond JSON or raw V2000 SDF. Optional Lipinski and Veber assessments are calculated from retrieved properties rather than predicted by a trained model.

The documented deployment choices are a local STDIO process, a self-hosted Streamable HTTP server and a public hosted endpoint. Calls are read-only, with request queuing, retries and cancellation handling. Structured statuses distinguish missing records from absent data, while pagination and truncation fields help clients handle bounded responses.

Coverage remains dependent on PubChem records: some compounds lack GHS information or computed 3D coordinates. Search and enrichment operations have explicit batch and pagination limits; descriptions and classification are fetched for only the first 10 CIDs in a batch. Some MCP clients expose tools but not resources, so resource access should not be assumed when planning an integration.

Key Features

  • Search compounds by name, SMILES, InChIKey, Hill formula, substructure, superstructure or 2D Tanimoto similarity, with optional property enrichment and unresolved-identifier reporting.
  • Retrieve selected physicochemical properties for up to 100 CIDs, with optional synonyms, descriptions, pharmacological classification and property-derived Lipinski/Veber assessments.
  • Fetch 2D PNG structure diagrams and 3D conformers as atom-and-bond JSON or V2000 SDF, with explicit preview caps and a distinct error for unavailable 3D structures.
  • Retrieve attributed GHS safety records for batches of up to 25 CIDs, distinguishing missing classifications from nonexistent compound records.
  • Explore bioactivity by outcome or molecular target, search assays by target identifiers, and retrieve entity summaries, interactions and external database cross-references.
  • Expose read-only tools and URI-templated resources over STDIO or Streamable HTTP, with queued requests, retry handling, cancellation propagation and disclosed partial failures.

Use Cases

  • Suggested evaluation: resolve a mixed list of compound names and SMILES to CIDs, inspect unresolved inputs, and enrich matched records with selected physicochemical properties.
  • Suggested evaluation: gather GHS classifications and literature or patent cross-references for a compound inventory, retaining source attribution and missing-data statuses.
  • Suggested evaluation: explore a target's PubChem bioassays and filter a compound's recorded activity profile by outcome or target before reviewing individual assay summaries.
  • Suggested evaluation: retrieve available 3D conformers for a structure-processing workflow, checking missing-coordinate errors and preview truncation before downstream use.

How to Use

  1. Read the project README to choose between hosted Streamable HTTP, local STDIO and self-hosted HTTP. Check the documented runtime prerequisites if running locally.
  2. For hosted access, configure your MCP client with the documented Streamable HTTP endpoint, https://pubchem.caseyjhand.com/mcp. For local deployment, follow the README's package or container configuration rather than substituting undocumented options.
  3. Begin an evaluation with pubchem_search_compounds and a small set of known names, SMILES or InChIKey inputs. Supply the fields required by the selected strategy and inspect unresolvedIdentifiers and match counts.
  4. Pass resolved CIDs to pubchem_get_compound_details, selecting properties needed for your workflow. Add safety, cross-reference or structure tools as appropriate; distinguish absent data from missing records.
  5. For biological exploration, use pubchem_search_assays with a supported target identifier and inspect relevant summaries or bioactivity pages. Follow offsets and truncation indicators, then compare representative outputs with the upstream records before adopting the integration.

Related resources

PubChem MCP (PhelanShao) is a Python MCP server that lets AI clients retrieve PubChem compound properties and structures by name or CID, with JSON, CSV, XYZ and downloadable structure-file outputs.

Open sourcePython

Cheminformatics · Scientific Data

PubChem PUG REST is an HTTP API for retrieving selected chemical records, properties and assay information, with documented identifier inputs, structure searches and operation-specific output formats.

Cheminformatics · Scientific Data

PubChemPy

Open Source

PubChemPy is a Python wrapper for the PubChem PUG REST API, supporting chemical searches, compound-property retrieval, standardization, file-format conversion and depiction.

Open sourcePython

Cheminformatics · Scientific Data

Scientific Agent Skills provides procedural guidance for AI agents working with scientific packages, databases and research workflows, including cheminformatics, spectroscopy and materials analysis.

Open sourcePython

Cheminformatics · Scientific Data

Benchling MCP (longevity-genie) is a Python MCP server that connects AI clients to Benchling notebook entries, biological sequences, projects, and entity search using API credentials.

Open sourcePython

Lab Automation · Scientific Data

ChEMBL MCP (cyanheads) connects MCP clients to ChEMBL compound, target, bioactivity and drug records, with structure searches and optional DuckDB analysis of larger activity sets.

Open sourceTypeScript

Drug Discovery · Scientific Data

Works with

Used in Recipes

Official

Chemistry Research with Claude + PubChem

Resolve chemical names to PubChem identifiers and retrieve traceable compound properties from Claude Desktop through a locally installed PubChem MCP server.

PubChem MCP (cyanheads) + PubChem PUG REST

Analyze Molecules

Level: Intermediate Cost: Mixed Privacy: Cloud ~30 min

Claude Desktop

View Setup →

Related guides

Overview

How to Build a Chemistry AI Agent Stack

Design a chemistry AI agent stack by separating language models, orchestration, Skills, MCP interfaces, APIs, libraries and data. Follow proposed molecular and materials workflows, define input/output contracts, and plan permissions, evaluation and reproducible deployment.

APIs

Connecting PubChem data to an AI agent

Plan a traceable PubChem lookup workflow using explicit compound identifiers, bounded tool calls, preserved provenance and checks for ambiguity, missing data and partial results.