Overview
ChEMBL Webresource Client is a Python library developed and supported by the ChEMBL group for accessing ChEMBL data and cheminformatics tools. It serves as a programmatic access layer to ChEMBL web services, rather than a standalone database or predictive model. The documented interface handles HTTPS communication so users can work with service resources without writing SQL or managing REST requests directly.
Its query design follows Django’s QuerySet interface. Inputs include endpoint-specific field filters and optional lists of fields to return; outputs are service results that the client retrieves when needed and caches on the local file system. Supported lookups include exact and case-insensitive matching, membership, numerical comparisons, ranges, null checks, regular expressions and search. Lazy evaluation postpones a data request until a value is required, making the library relevant to Python workflows that build queries before consuming records.
The only method lets users restrict returned fields, but those fields must exist on the selected endpoint. Nested field selections are not preserved at that granularity: selecting molecule_properties__alogp is treated as selecting molecule_properties. The README also states that only does not provide SQL join optimization for many-to-many relationships.
Configuration covers local caching, cache expiry, retry counts, request concurrency and timeouts. The documented FAST_SAVE setting carries a possibility of cache data loss. These controls are useful considerations when planning repeated retrieval workflows, but the source excerpts do not establish measured performance, endpoint-specific output schemas or tested scientific outcomes.
Key Features
- Provides an official Python interface to ChEMBL data and cheminformatics tools while handling HTTPS communication.
- Supports QuerySet-style filters including exact matching, case-insensitive text matching, membership, comparisons, ranges, null checks, regex and search.
- Evaluates queries lazily, requesting data only when a value is needed.
- Caches retrieved results locally, with configurable cache enablement, expiry and SQLite cache filename.
- Restricts returned fields through `only`, subject to endpoint field availability and documented nested-field limitations.
- Exposes settings for timeouts, HTTP retries, request concurrency and cache-saving behavior.
Use Cases
- Suggested evaluation: build filtered ChEMBL retrieval workflows in Python for drug-discovery data preparation without implementing REST request handling.
- Suggested evaluation: retrieve selected record fields for downstream analysis, checking whether nested-field behavior provides the required level of detail.
- Suggested evaluation: assess local caching and lazy retrieval for repeated exploratory queries using representative workloads and chosen cache settings.
How to Use
- Read the official client README to understand the query interface, installation guidance and configuration options before integrating the library into a Python workflow.
- Follow the README’s package installation instructions. Use its linked example notebook as a starting reference; availability has not been established here.
- Consult the ChEMBL API documentation to identify a suitable endpoint and its fields. Choose documented lookup types that match the records you intend to retrieve.
- If reducing returned fields, apply
onlyto fields available on that endpoint. Check its nested-field limitation before relying on a narrowly selected property. - Configure settings before using the client, as directed in the README. Review cache expiry, retries, concurrency and timeout needs; consider the stated data-loss risk of
FAST_SAVE. - Evaluate a small representative query first, inspecting returned fields and retrieval behavior before expanding the workflow. This is a proposed evaluation step, not a report of completed testing.