Overview
PubChem PUG REST provides programmatic access to selected PubChem information through HTTP requests. It supports focused retrieval in scripts, web applications and third-party tools, rather than requiring an installable client package. The official specification describes URLs organized into input, operation and output components. This entry concerns the hosted API itself, not a wrapper or third-party MCP server.
Compound inputs include CID, name, SMILES, InChI, InChIKey and molecular formula, alongside structure-search inputs. Substance and assay domains are also documented. Available operations include retrieving full records, synonyms, identifiers, compound properties and assay summaries, plus similarity and identity searches. Outputs depend on the operation and can include JSON, XML, SDF, PNG, TXT or CSV. These choices allow a workflow to request a specific property table or structure record instead of treating every request as an identical record export.
The tutorial positions PUG REST for focused requests, not millions of individual calls. Clients should respect published request limits, account for dynamic throttling and use bounded retries; bulk retrieval requires suitable bulk-download workflows. Special URL characters need encoding, while InChI and SDF inputs use POST. The programmatic-access overview distinguishes its short synchronous requests from PUG View’s complete summary reports and third-party annotations. Retrieving data does not experimentally validate bioactivity or establish reuse rights for every annotation. The supplied evidence establishes neither a service code repository nor a code license.
Key Features
- HTTP request paths separate the input specification, requested operation and output format.
- Compound queries accept CID, name, SMILES, InChI, InChIKey and formula inputs; substance and assay domains are also documented.
- Retrieval operations cover full records, synonyms, identifiers, compound properties and assay summaries.
- Structure-oriented operations include similarity and identity searches.
- Operation-dependent outputs include JSON, XML, SDF, PNG, TXT and CSV; InChI and SDF inputs use POST.
- Focused synchronous access is documented with request limits and dynamic throttling, rather than unrestricted bulk per-record retrieval.
Use Cases
- Suggested application: enrich a curated compound list with documented properties while preserving the PubChem identifier and retrieval date for each result.
- Suggested application: obtain SDF structure records for downstream inspection or cheminformatics processing, evaluating their suitability separately.
- Suggested application: retrieve synonyms and identifiers to support a chemical-name reconciliation workflow; assess ambiguous matches before accepting them.
- Intended evaluation: assess whether documented similarity searches or assay summaries meet a research workflow’s information needs without treating retrieved results as experimental validation.
How to Use
- Consult the PUG REST specification and choose the appropriate domain, input identifier, operation and output format. Prefer a stable record identifier when available.
- Inspect the documented aspirin SDF example to understand a structure-record request before adapting it to your chosen compound.
- For focused property retrieval, examine the MolecularFormula and InChIKey JSON example. Confirm that the returned fields meet your downstream requirements.
- Follow the specification’s input rules: encode special URL characters and use POST for InChI or SDF inputs. Inspect HTTP status and returned data, retaining the source identifier and retrieval date.
- Apply rate limiting and bounded retries using the official tutorial. Plan bulk retrieval separately rather than issuing unrestricted parallel per-record requests.
- Review the programmatic-access overview when choosing between focused PUG REST retrieval and PUG View reports. Evaluate data interpretation and annotation reuse requirements separately from API access.