Skip to content

PubChemPy

Matt Swain

PubChemPy is a Python wrapper for the PubChem PUG REST API, supporting chemical searches, compound-property retrieval, standardization, file-format conversion and depiction.

Catalog updated ·

Overview

PubChemPy provides a Python interface to the PubChem PUG REST API. Its role is to connect Python chemical-data workflows with PubChem searches and compound information, rather than to supply a separate chemical database or a predictive model. The README describes searches by chemical name, substructure and similarity, alongside chemical standardization, file-format conversion, depiction and property retrieval.

The supplied examples show two entry points: retrieving a compound using a PubChem Compound Identifier (CID), and searching by name. A CID lookup produces a compound object whose attributes expose information such as SMILES and an IUPAC name. The name-search example returns a collection of compound results and reads the first result’s SMILES and molecular weight. These examples illustrate how identifiers or names can become structured information accessible within Python.

In a data-preparation workflow, PubChemPy can serve as the retrieval layer between a list of chemical queries and downstream processing. Suggested evaluations include enriching records with PubChem properties or checking whether name-based results identify the intended compounds. Those applications require checking returned records; the README’s use of the first search result is an example, not evidence that every name resolves uniquely or appropriately.

The project metadata declares Python >=3.10 and no required package dependencies. The source excerpts do not establish service availability, retrieval throughput, database coverage or scientific validation of returned records. They also list broader operations without showing their detailed inputs and outputs, so the linked guide and API reference are the appropriate next resources for planning those workflows.

Key Features

  • Searches PubChem chemicals by name, substructure and similarity.
  • Retrieves compound records by CID through `Compound.from_cid`.
  • Exposes compound attributes including SMILES, IUPAC name and molecular weight in the documented examples.
  • Supports chemical standardization through its PubChem interface.
  • Provides chemical file-format conversion and depiction capabilities described in the README.

Use Cases

  • Suggested evaluation: enrich a chemical inventory containing PubChem CIDs with SMILES and IUPAC names, checking retrieved records against the intended substances.
  • Suggested evaluation: resolve chemical names to candidate PubChem records and inspect result selection before using molecular weights in downstream data preparation.
  • Suggested evaluation: assess the documented substructure or similarity searches for assembling candidate compound sets for a cheminformatics workflow.
  • Suggested evaluation: examine standardization, format conversion or depiction workflows in the API documentation using representative chemical inputs.

How to Use

  1. Start with the repository README to choose between CID retrieval, name search or another described operation. Define which compound attributes your workflow needs before selecting an example.
  2. Consult the installation guide and check the declared Python >=3.10 requirement. The README supplies pip and conda installation options; use those documented instructions rather than adding unsupported flags.
  3. Follow the getting-started guide. For an initial evaluation, reproduce either the README’s CID 1423 lookup or its Aspirin name search and inspect the returned attributes.
  4. Use the API reference to investigate inputs and outputs for substructure searches, similarity searches, standardization, conversion or depiction. Detailed procedures for these operations are not present in the source excerpts.
  5. Evaluate representative records before integrating retrieval into a larger workflow. Check compound identity and required fields, record unresolved queries, and consult the issue tracker when investigating unexpected behavior.

Related resources

PubChem MCP (cyanheads) connects MCP clients to PubChem compound and bioassay data, with identifier and structure searches, property retrieval, safety records, cross-references and 3D conformers.

Open sourceTypeScript

Cheminformatics · Scientific Data

PubChem MCP (PhelanShao) is a Python MCP server that lets AI clients retrieve PubChem compound properties and structures by name or CID, with JSON, CSV, XYZ and downloadable structure-file outputs.

Open sourcePython

Cheminformatics · Scientific Data

PubChem PUG REST is an HTTP API for retrieving selected chemical records, properties and assay information, with documented identifier inputs, structure searches and operation-specific output formats.

Cheminformatics · Scientific Data

Scientific Agent Skills provides procedural guidance for AI agents working with scientific packages, databases and research workflows, including cheminformatics, spectroscopy and materials analysis.

Open sourcePython

Cheminformatics · Scientific Data

Benchling MCP (longevity-genie) is a Python MCP server that connects AI clients to Benchling notebook entries, biological sequences, projects, and entity search using API credentials.

Open sourcePython

Lab Automation · Scientific Data

ChEMBL MCP (cyanheads) connects MCP clients to ChEMBL compound, target, bioactivity and drug records, with structure searches and optional DuckDB analysis of larger activity sets.

Open sourceTypeScript

Drug Discovery · Scientific Data

Works with

Used in Recipes

Official

Chemical Data Cleaning Stack

Validate and canonicalize SMILES with RDKit without dropping original rows; flag duplicates and optionally enrich valid entries with PubChem identifiers.

RDKit + PubChemPy

Chemical Data

Level: Intermediate Cost: Free Privacy: Depends on configuration ~30 min

Python

View Setup →

Related guides

Overview

How to Build a Chemistry AI Agent Stack

Design a chemistry AI agent stack by separating language models, orchestration, Skills, MCP interfaces, APIs, libraries and data. Follow proposed molecular and materials workflows, define input/output contracts, and plan permissions, evaluation and reproducible deployment.

APIs

Connecting PubChem data to an AI agent

Plan a traceable PubChem lookup workflow using explicit compound identifiers, bounded tool calls, preserved provenance and checks for ambiguity, missing data and partial results.