Why APIs still belong in a chemistry AI stack
A chemistry assistant should retrieve scientific records rather than reconstruct them from model memory. APIs give developers a defined route to provider identifiers, selected fields, and structured responses. The application can then check those responses before presenting a scientific answer.
PubChem PUG REST serves focused chemical-record requests. RCSB PDB Data API supplies metadata about identified structures and their components. Neither turns retrieval into experimental validation, structure prediction, or evidence of drug efficacy.
Start with direct API calls when you need predictable requests and inspectable responses. Consider an MCP wrapper when an assistant needs to choose among exposed tools. The API remains the upstream data interface; the wrapper adds another contract and failure boundary. The API, MCP, Skills, and agents guide explains these roles, while MCP in chemistry provides broader integration context.
Related guides: RDKit AI Workflow · Protein and Structural Biology Workflow · Chemistry AI Agent Stack.
Define the interface contract before choosing tools
Specify the scientific task first: resolving a compound identifier, retrieving a property, inspecting a polymer sequence, or following a database cross-reference. Write down what the result may support—and what it must not imply.
A proposed minimum contract contains:
- Inputs: identifier namespace and value, requested fields, output format, and filters.
- Outputs: provider identifiers, returned representations, values and units where applicable, available source references, retrieval time, and result status.
- Missingness: distinct handling for absent fields, explicit nulls, empty results, and numerical zero.
- Identity: provider identifiers retained alongside internal identifiers.
- Completeness: whether the response is complete, partial, paginated, or cached.
Decide whether the application requires an exact chemical entity or a broader grouping of related forms. Salt components, charge, and stereochemistry should not disappear through an undocumented normalization step. Likewise, an entry identifier, entity identifier, and chain identifier should not become interchangeable merely because they appear in the same structure record.
Choose the service by its role
These resources occupy different stages. The following table compares documented capabilities, not tested interoperability.
| Resource | Source-described role | Suitable input or output | Integration boundary |
|---|---|---|---|
| PubChem PUG REST | Hosted chemical-data API | CID, name, or structure input; selected records and properties | Not a complete PUG View summary report |
| RCSB PDB Data API | Hosted structure-metadata API | Identified entries, entities, instances, assemblies, and chemical components; JSON | Discovery uses a separate Search API; metadata is not a coordinate-file download |
| UniProt MCP (cyanheads) | Third-party protein-data wrapper | Protein queries, accessions, mapping jobs, annotations, and sequences | Mapping can require polling and page continuation |
| RCSB MCP (cnyambura) | Third-party retrieval and download wrapper | PDB entry and polymer lookups; downloaded-file status and local path | Organism-search tool returns instructions, not retrieved matches |
| RDKit MCP Server (TandemAI) | Software interface to RDKit functions | Exposed RDKit tool results and client responses | Individual molecular input schemas and comprehensive tool coverage are not established here |
The separate rcsb-api Python client provides access to RCSB Search and Data APIs; it is not the hosted Data API implementation. Similarly, an MCP server is executable software, not merely a Skill instruction document. Loading workflow instructions does not install dependencies or establish host compatibility. Use the MCP selection guide and agents and Skills guide when choosing that additional layer.
Start with focused PubChem REST requests
PUG REST paths divide a request into input, operation, and output. Compound inputs include CID, name, SMILES, InChI, InChIKey, and molecular formula. Documented operations include records, synonyms, identifiers, properties, assay summaries, and structural identity or similarity searches. Output formats depend on the operation.
Two official examples make useful starting points:
- Send an HTTP GET to the CID 2244 property request to request
MolecularFormulaandInChIKeyas JSON. - Send an HTTP GET to the CID 2244 SDF request to request the documented aspirin structure record.
These examples illustrate different outputs for the same identifier; they are not execution results reported by this guide. Inspect HTTP status, response format, returned identifiers, and requested-field presence before accepting either response.
Encode special URL characters rather than inserting a molecular representation unmodified into a URL. The specification directs InChI and SDF inputs through POST. For structure searching, distinguish an identity criterion from a similarity criterion: candidate similarity does not authorize merging records as the same compound. Confirm the specific operation and parameters in the specification before implementing it.
PUG REST returns selected information through focused synchronous requests. PUG View supplies fuller summary reports and third-party annotations. A property lookup therefore should not be presented as the complete PubChem report. For further application planning, see the PubChem agent integration guide.
Retrieve RCSB metadata at the correct hierarchy level
RCSB REST supports GET requests with resource-specific paths and JSON responses. Its official examples include:
- Entry
4HHB: entry-level metadata, including experimental details. - Polymer entity
4HHB/1: information about a chemically unique polymer. - Polymer instance
4HHB/A: information about a particular copy of that polymer.
In the instance endpoint, the chain identifier corresponds to _label_asym_id in PDBx/mmCIF. Preserve that namespace when connecting metadata to structural files.
REST returns a fixed object representation. GraphQL permits selected fields and traversal across connected hierarchy levels, including multiple identified objects. Choose REST for a straightforward object lookup; consider GraphQL when the application needs a narrow selection across relationships. Both interfaces query the same underlying data.
Keep discovery separate from retrieval. The Search API finds lists of matching identifiers; the Data API retrieves information about supplied identifiers. GraphQL does not provide an all-objects query. Archive-wide workflows start with the separate Holdings service and retrieve metadata in batches.
The Data API covers commonly used annotations, not every PDBx/mmCIF item. Coordinate files are also a separate retrieval concern. The RCSB MCP wrapper documents file downloads, but a successful metadata request alone does not establish that a coordinate file has been obtained.
For discovery, follow the Search API documentation: choose an attribute, sequence or structure-search service, declare return_type, and bound pagination in request_options. Retain the query and returned identifier namespace; then retrieve metadata through Data API. Search matches remain candidates requiring scientific interpretation.
Proposed example: a compound and structure evidence card
Consider a proposed internal application with two independent inputs: PubChem CID 2244 and PDB entry 4HHB. Their use together illustrates interface design; it does not assert a compound–protein relationship.
- Request the documented PubChem property JSON and, if needed, SDF. Preserve the CID, returned descriptors, and retrieval time.
- Request RCSB entry metadata. Follow returned entity identifiers instead of assuming every entry uses the same entity numbering.
- Retrieve the selected polymer entity and instance when sequence-level and chain-specific annotations are both needed.
- Display separate compound and structure panels, each with provider attribution and status.
- Require an independently supported relationship before joining the panels into a biological claim.
Proposed output: identifiers, requested values, representations, experimental or model context where supplied, citations, completeness, and warnings. Decision points: Are identity checks satisfied? Is required context present? Is a cross-reference sufficient for the intended join? If not, return ambiguity or keep the records separate.
As a proposed evaluation, compare representative responses against official record pages and include invalid identifiers, missing fields, charged compounds, and stereochemical variants. No scientific testing or completed integration is claimed.
Add wrappers and agents without hiding upstream evidence
UniProt MCP can supply protein annotations and cross-references, with curation indicators and available PubMed/ECO evidence. A proposed extension would follow a returned PDB reference to RCSB metadata, then check organism, sequence, and identifier scope before accepting a join. Continue polling existing mapping tickets and fetch completed-result pages separately; do not repeatedly submit the same job.
RCSB MCP exposes entry and polymer retrieval, custom queries, and structure downloads. Its custom-endpoint examples use identifier punctuation that differs from official REST path examples. Check how the wrapper constructs requests rather than assuming those strings can be pasted directly into upstream REST URLs.
RDKit MCP offers a tool-listing utility and an evaluation suite. Inspect the actual inventory before proposing normalization or descriptor calculations. Full RDKit function coverage is an ambition, not demonstrated coverage. Keep calculated outputs separate from provider-returned fields.
For local UniProt deployment, resolve the prerequisite discrepancy: the README allows Bun v1.3.0 or higher, while package metadata requires Bun >=1.4.0. Configuration instructions are not evidence of tested compatibility with your chosen host.
Handle limits, failures, and changing responses
Configure shared request scheduling, timeouts, and bounded retries. PubChem documents request limits and dynamic throttling; use its tutorial rather than unrestricted parallel per-record calls, and choose bulk-download workflows for bulk data. RCSB recommends batching large requests and caching repeat retrievals. Neither guidance is a throughput guarantee.
Expose meaningful states such as found, not found, ambiguous, invalid input, access denied, partial, and temporarily unavailable. RCSB REST documents 404 for missing data or endpoints. GraphQL errors can accompany HTTP 200 and partial data, so inspect the response body before declaring success.
Retry transient failures only. Preserve pagination completeness, label cached responses with their observation time, and fail explicitly when stale data would invalidate the task. Keep credentials server-side where required and restrict user-triggered operations, especially custom endpoints and file destinations. Log diagnostic information without secrets or unnecessary sensitive inputs.
Preserve provenance and use a compact readiness checklist
Retain the provider, record identifier, result-determining request parameters, observation time, and available citation metadata. Label normalized structures, converted units, and calculated summaries as derived. Live records can change; retain permitted fixtures or snapshots and document historical reproducibility limits.
Retrieval does not prove bioactivity, validate a computed model, or establish annotation reuse rights. Check software, database, annotation, and hosted-service terms separately. Source availability is not itself a license determination. The dataset provenance guide expands this recordkeeping approach.
Before proceeding:
- Declare identifier namespaces and chemical identity rules.
- Confirm fields, units, missingness, and output formats.
- Separate Search API discovery from Data API retrieval.
- Inspect errors, partial results, and pagination.
- Establish throttling, retry, cache, and timeout policies.
- Inspect wrapper tool schemas and deployment requirements.
- Preserve citations and label transformations.
- Record unresolved scientific and reuse limitations.
Sources
The PubChem specification, tutorial, and programmatic-access overview support the request examples, focused-retrieval scope, and operational boundaries. RCSB Data API documentation supports hierarchy, REST examples, GraphQL handling, and the Search/Data distinction.
Wrapper capabilities come from the project READMEs for UniProt MCP, RCSB MCP, and RDKit MCP. The UniProt package metadata supports the runtime discrepancy discussed above.