Skip to content

ChEMBL Webresource Client

Michal Nowotka, Eloy Felix / ChEMBL group

Official Python client maintained by the ChEMBL group for querying ChEMBL data and cheminformatics services, with QuerySet-style filters, lazy retrieval and local result caching.

Catalog updated ·

Overview

ChEMBL Webresource Client is a Python library developed and supported by the ChEMBL group for accessing ChEMBL data and cheminformatics tools. It serves as a programmatic access layer to ChEMBL web services, rather than a standalone database or predictive model. The documented interface handles HTTPS communication so users can work with service resources without writing SQL or managing REST requests directly.

Its query design follows Django’s QuerySet interface. Inputs include endpoint-specific field filters and optional lists of fields to return; outputs are service results that the client retrieves when needed and caches on the local file system. Supported lookups include exact and case-insensitive matching, membership, numerical comparisons, ranges, null checks, regular expressions and search. Lazy evaluation postpones a data request until a value is required, making the library relevant to Python workflows that build queries before consuming records.

The only method lets users restrict returned fields, but those fields must exist on the selected endpoint. Nested field selections are not preserved at that granularity: selecting molecule_properties__alogp is treated as selecting molecule_properties. The README also states that only does not provide SQL join optimization for many-to-many relationships.

Configuration covers local caching, cache expiry, retry counts, request concurrency and timeouts. The documented FAST_SAVE setting carries a possibility of cache data loss. These controls are useful considerations when planning repeated retrieval workflows, but the source excerpts do not establish measured performance, endpoint-specific output schemas or tested scientific outcomes.

Key Features

  • Provides an official Python interface to ChEMBL data and cheminformatics tools while handling HTTPS communication.
  • Supports QuerySet-style filters including exact matching, case-insensitive text matching, membership, comparisons, ranges, null checks, regex and search.
  • Evaluates queries lazily, requesting data only when a value is needed.
  • Caches retrieved results locally, with configurable cache enablement, expiry and SQLite cache filename.
  • Restricts returned fields through `only`, subject to endpoint field availability and documented nested-field limitations.
  • Exposes settings for timeouts, HTTP retries, request concurrency and cache-saving behavior.

Use Cases

  • Suggested evaluation: build filtered ChEMBL retrieval workflows in Python for drug-discovery data preparation without implementing REST request handling.
  • Suggested evaluation: retrieve selected record fields for downstream analysis, checking whether nested-field behavior provides the required level of detail.
  • Suggested evaluation: assess local caching and lazy retrieval for repeated exploratory queries using representative workloads and chosen cache settings.

How to Use

  1. Read the official client README to understand the query interface, installation guidance and configuration options before integrating the library into a Python workflow.
  2. Follow the README’s package installation instructions. Use its linked example notebook as a starting reference; availability has not been established here.
  3. Consult the ChEMBL API documentation to identify a suitable endpoint and its fields. Choose documented lookup types that match the records you intend to retrieve.
  4. If reducing returned fields, apply only to fields available on that endpoint. Check its nested-field limitation before relying on a narrowly selected property.
  5. Configure settings before using the client, as directed in the README. Review cache expiry, retries, concurrency and timeout needs; consider the stated data-loss risk of FAST_SAVE.
  6. Evaluate a small representative query first, inspecting returned fields and retrieval behavior before expanding the workflow. This is a proposed evaluation step, not a report of completed testing.

Related resources

ChEMBL MCP (cyanheads) connects MCP clients to ChEMBL compound, target, bioactivity and drug records, with structure searches and optional DuckDB analysis of larger activity sets.

Open sourceTypeScript

Drug Discovery · Scientific Data

PocketScout MCP equips AI assistants to assess known protein binding sites using public structural, bioactivity, conservation, variant, and literature data before target selection or binder design.

Open sourcePython

Drug Discovery · Scientific Data

RCSB MCP (cnyambura) is a third-party MCP server that exposes RCSB PDB entry and polymer metadata, structure downloads, and custom Data API queries to an MCP client.

Python

Drug Discovery · Scientific Data

RCSB PDB Data API provides REST and GraphQL access to structure metadata, molecular sequences, chemical descriptors, and annotations for PDB entries and selected computed structure models.

Drug Discovery · Scientific Data

UniProt MCP (cyanheads) is a third-party MCP server that connects AI clients to UniProt protein search, annotations, identifier mapping, proteomes, taxonomy, and amino-acid sequences.

Open sourceTypeScript

Drug Discovery · Scientific Data

AIDDISON Explorer is a hosted drug-discovery platform that generates and ranks molecular candidates against target profiles, design constraints, predicted properties and synthetic feasibility.

Molecular Generation · Drug Discovery

Related guides

Workflows

Protein & Structural Biology AI Workflow

Build an evidence-led protein workflow from UniProt identity and sequences to RCSB structure metadata, observed ligand contacts, and proposed downstream candidate evaluation—with explicit checks for species, isoforms, chains, missing data, and experimental validation.