Skip to content

RCSB PDB Data API

RCSB PDB

RCSB PDB Data API provides REST and GraphQL access to structure metadata, molecular sequences, chemical descriptors, and annotations for PDB entries and selected computed structure models.

Catalog updated ·

Overview

RCSB PDB Data API is a metadata service for experimental structures in the PDB, selected Computed Structure Models (CSMs), and integrative structures. It supplies molecule names, sequences, experimental details, and related annotations for structural-biology workflows. Its role is to retrieve documented information about identified structures and their components, rather than predict structures or perform molecular simulations. The official documentation describes both interfaces and their shared underlying data.

Requests use entry identifiers or identifiers for entities, chains, assemblies, and chemical components; responses are JSON. REST retrieves a fixed object representation through resource-specific endpoints. GraphQL lets users select fields and follow relationships between hierarchy levels in one request, including requests for multiple identified objects. The hierarchy distinguishes chemically unique molecules from their individual chain instances, preserving the difference between sequence-level information and instance-specific structural annotations.

Available information includes citations, taxonomy, sequence features, chemical descriptors, and reference-database identifiers. Separate examples cover computed-model quality annotations and the composition and input datasets of integrative structures. Coverage is limited to commonly used annotations, not every item in the PDBx/mmCIF dictionary. GraphQL requires explicit identifiers rather than offering an all-objects query; archive-wide retrieval therefore starts with the separate Holdings service and proceeds in batches. Applications must also inspect GraphQL response errors, not just HTTP status. The documented Python package, rcsb-api, is a separate client for the Search and Data APIs, not the Data API service itself.

Key Features

  • REST GET endpoints return JSON objects for entries, molecular entities, entity instances, assemblies, and chemical components.
  • GraphQL supports field selection, single or multiple object identifiers, and traversal between connected levels of the structure hierarchy.
  • Polymer annotations include sequences, taxonomy, sequence-cluster membership, positional features, and reference-sequence cross-references.
  • Chemical-component records expose descriptors such as SMILES and InChI, chemical formulas, and systematic chemical names; branched-entity examples cover carbohydrate descriptors.
  • Entry-level retrieval includes experimental details, primary citations, release information, and links to constituent entities.
  • Documented queries cover global quality-assessment metrics for computed models and composition, input datasets, and source databases for integrative structures.

Use Cases

  • Suggested application: enrich a known set of PDB identifiers with experimental methods, release dates, and primary citations for a structure-selection report.
  • Suggested application: assemble protein sequence, taxonomy, domain, and chain-level feature annotations for downstream structural-biology analysis.
  • Suggested application: retrieve ligand chemical descriptors and associated structure metadata as inputs to a drug-discovery data-curation workflow, without treating them as evidence of binding activity.
  • Suggested evaluation: build a batched archive-wide metadata collection and check field coverage, missing cross-references, and distinctions between experimental, computed, and integrative records.

How to Use

  1. Read the data organization guide and identify the appropriate object level. Keep entity identifiers separate from chain identifiers; polymer instances use label_asym_id.
  2. Choose REST for a fixed object response or GraphQL for selected, related fields. Consult the REST reference or explore the schema in GraphiQL.
  3. Start with a documented example, such as entry 4HHB, and check its returned JSON against the fields needed for your workflow. Use the Data Attributes page to verify annotation availability.
  4. Implement response checks using the API guidance: distinguish REST status failures from GraphQL errors included in the JSON response, even when HTTP status is successful.
  5. For archive-wide retrieval, obtain identifiers from the current holdings endpoint, then follow the batching and caching guidance. If a Python client is useful, consult the separate rcsb-api documentation.

Related resources

ChEMBL MCP (cyanheads) connects MCP clients to ChEMBL compound, target, bioactivity and drug records, with structure searches and optional DuckDB analysis of larger activity sets.

Open sourceTypeScript

Drug Discovery · Scientific Data

Official Python client maintained by the ChEMBL group for querying ChEMBL data and cheminformatics services, with QuerySet-style filters, lazy retrieval and local result caching.

Open sourcePython

Drug Discovery · Scientific Data

PocketScout MCP equips AI assistants to assess known protein binding sites using public structural, bioactivity, conservation, variant, and literature data before target selection or binder design.

Open sourcePython

Drug Discovery · Scientific Data

RCSB MCP (cnyambura) is a third-party MCP server that exposes RCSB PDB entry and polymer metadata, structure downloads, and custom Data API queries to an MCP client.

Python

Drug Discovery · Scientific Data

UniProt MCP (cyanheads) is a third-party MCP server that connects AI clients to UniProt protein search, annotations, identifier mapping, proteomes, taxonomy, and amino-acid sequences.

Open sourceTypeScript

Drug Discovery · Scientific Data

AIDDISON Explorer is a hosted drug-discovery platform that generates and ranks molecular candidates against target profiles, design constraints, predicted properties and synthetic feasibility.

Molecular Generation · Drug Discovery

Works with

Used in Recipes

Official

Protein Structure Research Stack

Combine RCSB structure records with UniProt protein annotations, inspect method/resolution and entity mappings, and download a traceable mmCIF file.

UniProt MCP (cyanheads) + RCSB MCP (cnyambura) + RCSB PDB Data API

Protein & Drug Research

Level: Intermediate Cost: Mixed Privacy: Cloud ~45 min

Claude Desktop · Python

View Setup →

Related guides

Overview

MCP vs Skills vs Agents vs APIs: What's the Difference?

Understand how APIs, MCP servers, Skills, and agents divide interface, procedural guidance, execution, and planning responsibilities in chemistry workflows—and how to choose or combine them without confusing access with scientific validity.

Workflows

Protein & Structural Biology AI Workflow

Build an evidence-led protein workflow from UniProt identity and sequences to RCSB structure metadata, observed ligand contacts, and proposed downstream candidate evaluation—with explicit checks for species, isoforms, chains, missing data, and experimental validation.