Skip to content

DeepChem

DeepChem contributors

DeepChem is an open-source Python toolchain for scientific deep learning in drug discovery, materials science, quantum chemistry and biology, with backend-specific dependencies and guided tutorials.

Catalog updated ·

Overview

DeepChem is a scientific machine-learning toolchain intended to support deep-learning work in drug discovery, materials science, quantum chemistry and biology. It is a library and collection of learning resources, rather than a single pretrained predictor or a hosted analysis service. The linked sources describe support for TensorFlow, PyTorch and JAX, allowing researchers to choose a framework according to the model they intend to use.

Its documented entry workflow begins with sequenced tutorials on molecular machine learning and computational biology, followed by additional examples. These tutorials are designed for Google Colab or local execution. For a new research problem, the project recommends adapting an existing tutorial or example incrementally. Scientific datasets serve as inputs to these workflows, but the source excerpts do not specify accepted file formats, representations or individual model output schemas. They do explicitly describe an integration with Weights & Biases for tracking model training and evaluation metrics.

Environment preparation is an important part of adoption. DeepChem has core scientific-computing dependencies, including NumPy, pandas, scikit-learn, SciPy and RDKit, plus optional requirements that depend on the selected functionality. Installation routes include pip, conda, Docker and source builds, with stable and nightly options. The README and package configuration differ in their Python compatibility declarations, so the excerpts do not establish one consistent supported range. They also provide no task-level performance results or evidence that a particular model will generalize to a new scientific dataset.

Key Features

  • Supports TensorFlow, PyTorch and JAX workflows through separately specified optional dependencies.
  • Provides a sequenced tutorial collection for molecular machine learning and computational biology, designed for Google Colab or local execution.
  • Offers additional examples as starting points for adapting DeepChem to new research problems.
  • Integrates with Weights & Biases to track model training and evaluation metrics.
  • Documents pip, conda, Docker and source installation routes, including stable and nightly builds.

Use Cases

  • Suggested evaluation: adapt a molecular machine-learning tutorial to a drug-discovery dataset and assess its suitability for the intended prediction task.
  • Suggested evaluation: investigate whether the documented models and examples meet the representation and output requirements of a materials-science or quantum-chemistry project.
  • Work through the sequenced tutorials to learn molecular machine learning and computational biology in Google Colab or a local environment.
  • Suggested evaluation: use the documented Weights & Biases integration to compare training and evaluation metrics across experimental workflows.

How to Use

  1. Start with the official documentation and identify a tutorial or example relevant to your scientific question. Check its actual input and output requirements rather than assuming a universal dataset interface.
  2. Consult the repository installation guidance to choose a stable, nightly, Docker or source installation route. Confirm requirements for the package you select, since the supplied Python declarations differ.
  3. Identify whether your chosen model needs TensorFlow, PyTorch or JAX. Review the soft requirements guidance for additional packages; GPU setup also depends on the selected framework.
  4. Work through a relevant tutorial in Google Colab or locally before changing the workflow. Record your environment and dataset assumptions.
  5. Adapt an existing example incrementally to your dataset. As an intended evaluation step, assess task-specific results and limitations before relying on predictions; optionally use the Weights & Biases integration to track metrics.

Related resources

AIDDISON Explorer is a hosted drug-discovery platform that generates and ranks molecular candidates against target profiles, design constraints, predicted properties and synthetic feasibility.

Molecular Generation · Drug Discovery

AutoDock Vina

Open Source

AutoDock Vina is an open-source molecular docking and virtual screening program with Vina and AutoDock4.2 scoring, multi-ligand docking, macrocycle support and Python 3 bindings.

Open sourcePython

Drug Discovery

Boltz

Model

Boltz is a biomolecular interaction model family. Boltz-2 provides workflows for complex structure and binding affinity prediction.

Open sourcePython

Drug Discovery

ChEMBL MCP (cyanheads) connects MCP clients to ChEMBL compound, target, bioactivity and drug records, with structure searches and optional DuckDB analysis of larger activity sets.

Open sourceTypeScript

Drug Discovery · Scientific Data

Official Python client maintained by the ChEMBL group for querying ChEMBL data and cheminformatics services, with QuerySet-style filters, lazy retrieval and local result caching.

Open sourcePython

Drug Discovery · Scientific Data

Chemistry42

Platforms

Chemistry42 is a small-molecule discovery platform combining generative design, retrosynthesis, ADMET and selectivity prediction, and physics-based prioritization for hit identification and lead optimization.

Molecular Generation · Drug Discovery

Related guides

Models

Choosing between Chemprop and DeepChem

Compare Chemprop’s molecular property prediction focus with DeepChem’s broader scientific machine-learning scope, then plan a fair evaluation using shared data, explicit inputs and outputs, and reproducible decision criteria.

Deployment

Local Deployment Checklist for Chemistry AI

Plan an isolated, reproducible local environment for RDKit, Chemprop or DeepChem. Check documented dependencies, data movement, example outputs and recovery procedures before using sensitive chemistry data.

Models

Evaluating Molecular Property Prediction Models

A practical guide to choosing molecular property workflows, checking representations and evaluation evidence, and planning a local assessment with explicit domain boundaries and component-level licence checks.