Overview
RCSB PDB Data API is a metadata service for experimental structures in the PDB, selected Computed Structure Models (CSMs), and integrative structures. It supplies molecule names, sequences, experimental details, and related annotations for structural-biology workflows. Its role is to retrieve documented information about identified structures and their components, rather than predict structures or perform molecular simulations. The official documentation describes both interfaces and their shared underlying data.
Requests use entry identifiers or identifiers for entities, chains, assemblies, and chemical components; responses are JSON. REST retrieves a fixed object representation through resource-specific endpoints. GraphQL lets users select fields and follow relationships between hierarchy levels in one request, including requests for multiple identified objects. The hierarchy distinguishes chemically unique molecules from their individual chain instances, preserving the difference between sequence-level information and instance-specific structural annotations.
Available information includes citations, taxonomy, sequence features, chemical descriptors, and reference-database identifiers. Separate examples cover computed-model quality annotations and the composition and input datasets of integrative structures. Coverage is limited to commonly used annotations, not every item in the PDBx/mmCIF dictionary. GraphQL requires explicit identifiers rather than offering an all-objects query; archive-wide retrieval therefore starts with the separate Holdings service and proceeds in batches. Applications must also inspect GraphQL response errors, not just HTTP status. The documented Python package, rcsb-api, is a separate client for the Search and Data APIs, not the Data API service itself.
Key Features
- REST GET endpoints return JSON objects for entries, molecular entities, entity instances, assemblies, and chemical components.
- GraphQL supports field selection, single or multiple object identifiers, and traversal between connected levels of the structure hierarchy.
- Polymer annotations include sequences, taxonomy, sequence-cluster membership, positional features, and reference-sequence cross-references.
- Chemical-component records expose descriptors such as SMILES and InChI, chemical formulas, and systematic chemical names; branched-entity examples cover carbohydrate descriptors.
- Entry-level retrieval includes experimental details, primary citations, release information, and links to constituent entities.
- Documented queries cover global quality-assessment metrics for computed models and composition, input datasets, and source databases for integrative structures.
Use Cases
- Suggested application: enrich a known set of PDB identifiers with experimental methods, release dates, and primary citations for a structure-selection report.
- Suggested application: assemble protein sequence, taxonomy, domain, and chain-level feature annotations for downstream structural-biology analysis.
- Suggested application: retrieve ligand chemical descriptors and associated structure metadata as inputs to a drug-discovery data-curation workflow, without treating them as evidence of binding activity.
- Suggested evaluation: build a batched archive-wide metadata collection and check field coverage, missing cross-references, and distinctions between experimental, computed, and integrative records.
How to Use
- Read the data organization guide and identify the appropriate object level. Keep entity identifiers separate from chain identifiers; polymer instances use label_asym_id.
- Choose REST for a fixed object response or GraphQL for selected, related fields. Consult the REST reference or explore the schema in GraphiQL.
- Start with a documented example, such as entry 4HHB, and check its returned JSON against the fields needed for your workflow. Use the Data Attributes page to verify annotation availability.
- Implement response checks using the API guidance: distinguish REST status failures from GraphQL errors included in the JSON response, even when HTTP status is successful.
- For archive-wide retrieval, obtain identifiers from the current holdings endpoint, then follow the batching and caching guidance. If a Python client is useful, consult the separate rcsb-api documentation.