Overview
ChEMBL MCP (cyanheads), packaged as @cyanheads/chembl-mcp-server, is a community integration for retrieving ChEMBL data through an MCP client. It wraps the upstream ChEMBL REST API rather than providing a separate scientific database or predictive model. Its workflow role is to connect compound discovery, protein-target identification and measured bioactivity retrieval within client-composed tool chains. The README documents local STDIO, Streamable HTTP and a public hosted endpoint.
Compound inputs include names, ChEMBL IDs, InChIKeys and SMILES for exact, similarity or substructure searches. Target searches accept UniProt accessions, gene symbols or free text, with organism and target-type filters. Returned molecule and target identifiers can feed bioactivity queries for a compound, a target or a specific compound–target pair. Outputs include molecular properties, activity measurements, drug mechanisms and indications, and assay provenance. Molecule and target records are also exposed as JSON MCP resources; the same data remains accessible through tools.
Bioactivity results can be filtered and ranked by pchembl_value, but the documentation restricts potency comparisons to a single standard_type. Missing potency remains null and can be retrieved separately through null_potency. Large activity sets can be staged in DuckDB-backed DataCanvas tables for read-only SQL analysis when that provider is enabled. Inline previews and staged datasets have separate limits, with truncation and row-count disclosures. Drug mechanism and indication lookups can return explicitly marked partial results. These controls help callers assess completeness and provenance; they do not establish scientific comparability without examining the underlying assays.
Key Features
- Search molecules by name, ChEMBL ID or InChIKey, or use SMILES for exact, Tanimoto similarity and substructure searches; results include molecular properties and maximum clinical phase.
- Resolve targets from UniProt accessions, gene symbols or text, returning target type, organism and component identifiers for subsequent bioactivity queries.
- Retrieve bioactivities by molecule, target or compound–target pair, with measurement-type, potency, assay-type and organism filters plus separate ranked and null-potency views.
- Retrieve drug mechanisms, molecular targets, action types, first-approval year and clinical indications, with completeness statuses for mechanism and indication lists.
- Inspect assay descriptions, types, organisms, targets and ChEMBL confidence scores behind individual bioactivity records.
- Stage large bioactivity sets in optional DuckDB-backed DataCanvas tables, inspect their schemas and run read-only SQL SELECT queries with disclosed result limits.
Use Cases
- Suggested evaluation: resolve a protein by UniProt accession, retrieve its IC50 measurements and inspect assay provenance before comparing candidate compounds.
- Suggested evaluation: search SMILES-based analogues and follow returned molecule identifiers into target-specific bioactivity queries.
- Suggested evaluation: assemble a drug background report linking documented mechanisms and indications to measured activities, retaining partial-result statuses.
- Suggested evaluation: group or deduplicate a staged activity dataset with read-only SQL, checking spill limits before treating aggregates as complete.
How to Use
- Read the project README and choose the documented hosted, local STDIO or self-hosted HTTP route. For local use, check its runtime prerequisites and client configuration examples.
- Configure an MCP client using those examples, or connect through Streamable HTTP to the public endpoint. The README states that upstream ChEMBL access requires no API key or account.
- Start a representative evaluation with
chembl_search_moleculesorchembl_search_targets. Supply a supported identifier or structure, and retain returned ChEMBL IDs for subsequent calls. - Call
chembl_get_bioactivitieswith a molecule ID, target ID or both. Select onestandard_typefor potency comparisons, inspect filters and counts, and examine missing-potency rows separately when relevant. - Follow assay IDs into
chembl_get_assay, or molecule IDs intochembl_get_drug_info. Check assay context and mechanism/indication completeness statuses before interpreting results. - For larger local datasets, enable
CANVAS_PROVIDER_TYPE=duckdbas documented. Inspect staged columns withchembl_dataframe_describe, then usechembl_dataframe_query; retain truncation disclosures and the README's ChEMBL attribution guidance.