Overview
UniProt MCP (cyanheads), distributed as @cyanheads/uniprot-mcp-server, provides an MCP interface to UniProt REST data. It is a third-party integration, not the UniProt database itself; the supplied sources do not establish upstream endorsement. Its role is to let an MCP client retrieve protein evidence and identifiers during research workflows, rather than predict protein properties or independently conduct research. The README describes local STDIO, self-hosted Streamable HTTP, and a public hosted endpoint.
Inputs include protein-function text, Lucene field queries, accessions, database identifiers, scientific organism names, and taxon or proteome IDs. Outputs include paginated search results, sectioned annotation records, identifier-mapping results, proteome metadata, taxonomy lineages, and canonical or optional isoform sequences. Search defaults to reviewed Swiss-Prot entries. Annotation outputs retain curation indicators and available PubMed/ECO evidence, allowing a client to distinguish annotation provenance. A supplied dossier prompt guides identifier resolution, annotation retrieval, and discovery of structure, bioactivity, and citation cross-references.
Several boundaries matter when integrating this server. Identifier mapping is asynchronous and can require polling a returned ticket, followed by separate result-page continuations. Large annotation records may return an outline that requires section-specific follow-up calls, while proteome protein listings are opt-in and capped. Missing upstream fields remain absent, and batch entry retrieval reports successes and failures separately. The two URI resources mirror entry and taxonomy tools; database-wide resource enumeration is not provided. Repository code is available under Apache-2.0, separately from UniProt data terms and operation of the hosted service.
Key Features
- Protein discovery through `uniprot_search_proteins`, with plain-text or Lucene queries, organism filtering, reviewed-entry defaults, optional facets, and forward cursor pagination.
- Sectioned annotation retrieval through `uniprot_get_entry`, supporting batches of up to 20 accessions, per-accession failures, field projection, and outline-based follow-up for oversized records.
- Asynchronous cross-database identifier mapping through `uniprot_map_ids`, with resumable job tickets, completed-page continuations, and taxon-based disambiguation of gene symbols.
- Proteome and taxonomy lookup by identifiers or scientific organism name, including proteome completeness metadata, lineage, optional immediate child taxa, and capped protein listings.
- Canonical FASTA sequence retrieval with length and parsed header, plus optional alternatively spliced isoform sequences.
- Entry and taxonomy URI resources, alongside the `uniprot_protein_dossier` prompt for guided annotation and cross-reference retrieval.
Use Cases
- Suggested evaluation: build a protein-target dossier combining documented function, localization, disease associations, variants, and PDB, ChEMBL, and PubMed cross-references.
- Suggested evaluation: reconcile Ensembl, RefSeq, gene-symbol, or other supported identifiers with UniProtKB accessions, explicitly checking species ambiguity and unmapped results.
- Suggested evaluation: assemble an organism-specific protein subset using taxonomy resolution, proteome metadata, and filtered protein listings.
- Suggested evaluation: retrieve canonical and isoform sequences alongside annotation provenance for downstream analysis, without treating retrieval as scientific validation.
How to Use
- Read the repository README and choose hosted Streamable HTTP or a local deployment. It supplies client configurations for STDIO and HTTP; UniProt REST access requires no API key.
- For the hosted route, configure your MCP client with https://uniprot.caseyjhand.com/mcp. For local setup, follow the README’s package or container configuration. Check package.json before selecting a runtime because its Bun requirement differs from the README prerequisite.
- Resolve the organism with
uniprot_get_taxonomy, then search by function or gene usinguniprot_search_proteins. Inspect reviewed status and annotation evidence rather than relying only on name matches. - Retrieve selected accessions with
uniprot_get_entry. Handle failed accessions individually and request specific sections when an outline is returned. For identifier mapping, retain tickets and continuations until the workflow finishes. - Retrieve sequences or proteome subsets as needed, or use
uniprot_protein_dossierfor a guided lookup. As an initial evaluation, compare retrieved annotations with UniProt and record unresolved identifiers and missing fields.