Start with an evidence question, not a design promise
A useful protein workflow asks a bounded question: which structural evidence supports investigating a particular target or binding region? The proposed sequence is: input sequence or identifier → protein identity and annotations → relevant structures → observed ligand contacts → candidate evaluation plan.
This guide combines source-described retrieval and triage capabilities with a proposed orchestration pattern. It does not describe a tested, ready-made integration. Use it for scientific-data assembly and early drug-discovery assessment, particularly when you need to explain why a structure or site deserves further investigation.
The deliverable should be an evidence packet, not an unsupported claim that a protein is druggable or a candidate will bind. Keep retrieved facts, software interpretations, and proposed research actions separate. For broader architectural context, see the chemistry AI stack guide and the MCP, Skills, agents, and APIs guide.
For component selection, read MCP for chemistry, AI Agents for Chemistry, and the broader drug-discovery workflow.
Choose the right component for each stage
These resources occupy different layers. The MCP servers are third-party interfaces or analysis tools; the RCSB PDB Data API is an upstream hosted metadata service; DrugAgent is a research framework described in a paper. None should be treated as interchangeable.
| Component | Source-described role | When to use it | Boundary to preserve |
|---|---|---|---|
| UniProt MCP (cyanheads) | Protein search, annotations, identifier mapping, taxonomy, proteomes, and sequences | Resolve identity and assemble a protein dossier | Retrieval is not protein-property prediction |
| RCSB MCP (cnyambura) | Entry and polymer metadata, structure downloads, custom Data API queries | Inspect identified PDB entries and obtain coordinate files | Its organism-search tool returns guidance, not retrieved search matches |
| RCSB PDB Data API | REST and GraphQL metadata for structures and their components | Retrieve precise entity, chain, assembly, and chemical-component relationships | Metadata requests do not perform docking or simulation |
| PocketScout MCP | Known-site triage using structural, bioactivity, conservation, variant, and literature evidence | Assess observed ligand-contact regions | No current computational pocket prediction; project metadata marks it Alpha |
| DrugAgent (Liu et al.) | Planner–Instructor collaboration for ML workflow programming and evaluation | Consider a separate downstream modelling investigation | No supplied evidence establishes integration with these MCP servers |
Choose explicit tools for narrow lookups and guided prompts for organizing a dossier. Prompts supply workflow instructions; they do not replace executable tools or dependencies.
Resolve protein identity before interpreting structure
Begin with the submitted identifier, scientific organism name, and sequence if available. Record whether the intended subject is a canonical protein, a particular isoform, or a modified experimental construct.
UniProt MCP provides uniprot_search_proteins, uniprot_map_ids, uniprot_get_entry, and uniprot_get_sequence. Search defaults to reviewed entries. Retain reviewed status, annotation provenance, protein-existence information, and available PubMed/ECO evidence rather than flattening all annotations into equally certain facts.
For a gene-symbol input, use organism filtering or taxon-based disambiguation. Retrieve the canonical sequence and optional isoform sequences, then propose a comparison against the submitted sequence. The supplied tool descriptions do not establish arbitrary sequence-similarity search: a sequence-only input needs a separately selected identification method or human resolution before accession-based retrieval.
The server documents removing an isoform suffix for entry and sequence calls. Preserve the original isoform identifier in the evidence packet; using a base accession must not silently change the biological subject. Pause if multiple proteins remain plausible or if the sequence relationship is unresolved.
Map structures at the entity and chain levels
Use available UniProt structure cross-references or PocketScout’s related-structure retrieval to establish candidate structure identifiers. Then inspect entry metadata and the relevant polymer entities. Do not select entity 1 merely because it is the RCSB MCP tool’s default.
The RCSB hierarchy distinguishes an entry from its chemically unique entities, individual entity instances or chains, biological assemblies, and chemical components. Sequence information shared across copies belongs at entity level; instance-specific structural annotations belong at chain level. The Data API uses label_asym_id for polymer-instance identifiers, so record the chain naming convention rather than assuming every displayed chain label is equivalent.
For each candidate structure, propose recording:
- Entry ID, target entity ID, chain instance, and assembly ID.
- Source organism, polymer sequence, reference-sequence cross-references, and any unresolved correspondence to the selected isoform.
- Structure category: experimental, computed model, or integrative structure.
- Experimental method and resolution where applicable, or the relevant model-quality and input-dataset annotations.
- Target-region coverage, coordinate availability, and ligand identifiers.
The Data API documents selected computed models and integrative structures, but that does not establish equivalent support in the older third-party RCSB MCP wrapper. Use the upstream interface where the wrapper’s documented scope is insufficient.
Move from structures to observed ligand contacts
RCSB MCP can download coordinate files, including mmCIF, and returns download status plus a local path. PocketScout documents coordinate analysis with gemmi to identify residues near co-crystallized ligands, classify sites, and consolidate observed pockets across structures by recurrence.
This distinction matters: chemical-component metadata identifies molecules, while coordinate analysis supplies spatial contact evidence. A ligand record alone is not a contact map. Likewise, a contact map is not a measured binding affinity or proof of functional modulation.
PocketScout also retrieves ligand bioactivity history, literature, known variants, and binding-residue conservation. Its detailed documentation describes local-context comparisons with mouse, rat, and cynomolgus macaque, not a full multiple sequence alignment. Preserve the raw information separately from each tool’s interpretation field.
For a proposed assessment, inspect whether the ligand and contacted chain correspond to the intended target. Treat pocket classification, size-based druggability assessment, and recurrence ranking as software-generated triage evidence, not experimental conclusions. If no suitable ligand-bearing structure exists, report that limitation; do not invent contact residues or substitute a predicted cavity without explicitly changing methods.
Connect MCP tools with explicit workflow gates
A proposed controller can pass verified identifiers between components and stop at unresolved decisions:
Identity gate → structure correspondence gate → coordinate and ligand gate → site-evidence assessment → human-approved downstream task.
UniProt MCP documents local STDIO, self-hosted Streamable HTTP, and a public hosted endpoint. PocketScout documents hosted and local operation. RCSB MCP documents standalone execution and Claude Desktop configuration. These descriptions are starting points for client evaluation, not proof that every current host or transport combination works.
Before local setup, resolve practical runtime differences. UniProt’s README prerequisite says Bun v1.3.0 or higher, whereas its package metadata requires Bun >=1.4.0; check the selected package’s requirements rather than relying on the lower threshold. RCSB MCP’s supplied project metadata requires Python >=3.13, while PocketScout requires Python >=3.11. Evaluate their environments separately instead of assuming a shared installation.
First propose checking tool discovery, input validation, output schemas, and file accessibility with a small retrieval task. A downloaded local path must be accessible to the intended analysis process. Loading a Skill or prompt supplies neither the server runtime nor coordinate-analysis dependencies. The chemistry MCP selection guide offers complementary selection criteria.
Proposed example: assess an EGFR structure without assuming the result
Proposed exercise, not a reported execution: use the EGFR/PDB 1M17 pairing referenced in PocketScout’s README to assemble a site-assessment packet. Start with the target name, intended species Homo sapiens, and the sequence or isoform used by the research team.
- Identity input: resolve the target in UniProt with the organism constraint. Output the chosen accession, canonical and relevant isoform sequences, and annotation provenance. Stop if the submitted sequence cannot be reconciled with the record.
- Structure input: retrieve metadata for
1M17and inspect its polymer entities and chains. Output a target-to-structure correspondence table. Continue only after checking that the relevant region and construct are suitable for the question. - Site input: request coordinate-based binding-site evidence. Output retrieved ligand identities and contact residues, retaining chain and residue-numbering context. If the response lacks contacts, record the gap rather than filling it from memory.
- Context input: use the verified accession and mapped residues for ligand history, conservation, variants, and literature retrieval. Output an evidence-backed list of reasons to investigate or defer each observed region.
- Decision output: propose one of three outcomes: investigate an observed site, seek additional structural evidence, or defer the target because identity or site evidence remains unresolved.
No particular ligand, contact residue, conservation result, or ranking is assumed here. A useful final packet includes raw responses, a correspondence table, unresolved questions, and an explicit validation plan.
Handle incomplete retrieval as a workflow state
UniProt mapping jobs may return a running-job ticket; continue polling that job rather than submitting it again. Completed mapping results can have separate page continuations. Search cursors, mapping continuations, and section outlines are not interchangeable.
For oversized annotation records, request the indicated sections. Inspect per-accession successes and failures in batches, and retain absent fields as absent. An incomplete response must not become a complete-sounding protein summary.
For RCSB requests, inspect REST status and GraphQL response errors: HTTP success alone does not establish a successful GraphQL query. GraphQL requires explicit object identifiers, and archive-wide retrieval starts separately with Holdings followed by batches. The Data API covers commonly used annotations, not every mmCIF dictionary item.
Propose bounded retries for transient failures, then mark the affected stage incomplete. Distinguish “no result retrieved,” “request failed,” and “evidence unavailable in the inspected sources.” Cache repeat retrievals with timestamps and retain query parameters; see the dataset provenance guide.
Evaluate candidates without overstating scientific evidence
The proposed candidate assessment should compare evidence completeness, target-region correspondence, observed site recurrence, ligand history, variant concerns, conservation, and contradictory literature. State reasons explicitly rather than inventing a universal numerical score.
A separate DrugAgent-style downstream task could receive a curated dataset, objective, constraints, starter files, and evaluator. Its paper describes planning and implementing ML workflows using performance or failure reports. This handoff is a proposed integration, not an existing UniProt/PDB/PocketScout connection. Supplied evidence does not establish usable installation instructions for DrugAgent, and the paper warns against direct pipeline deployment because of hallucination risks.
Structural confidence is not binding evidence; observed contacts are not selectivity evidence; conservation checks are not validation of an animal model. Before promoting a candidate, propose expert inspection and appropriate binding, functional, and selectivity experiments. No experimental validation is claimed for this workflow.
Reader checklist
- Identity, species, isoform, and submitted sequence are recorded separately.
- Entry, entity, chain, assembly, and residue numbering are unambiguous.
- Experimental structures, computed models, and integrative structures remain distinct.
- Ligand identities and contact evidence are retrieved, not assumed.
- Raw results, software interpretations, and research proposals are separated.
- Pagination, partial failures, missing fields, and GraphQL errors are handled.
- Client support, runtimes, dependencies, and file access are evaluated locally.
- Candidate decisions include unresolved evidence and experimental next steps.
Continue with these related guides: APIs for Chemistry AI Applications: Getting Started with PubChem and RCSB PDB.
Sources
The UniProt MCP README and package metadata support its retrieval functions and runtime caveat. The RCSB MCP README and project metadata describe the wrapper. RCSB PDB Data API documentation supports the hierarchy, interfaces, and error handling. The PocketScout README and project metadata describe known-site analysis and Alpha status. The DrugAgent paper supports its downstream programming role and limitations.