Skip to content

DeepPurpose

Kexin Huang, Tianfan Fu

DeepPurpose is a PyTorch-based toolkit for training and applying compound and protein prediction models, with workflows for drug-target interactions, molecular properties, repurposing and virtual screening.

Catalog updated ·

Overview

DeepPurpose is a molecular modeling toolkit centered on drug-target interaction prediction, with additional modules for compound properties, drug-drug interactions, protein-protein interactions and protein function. It provides a configurable training and inference framework rather than a single predictive model. Its documented workflow connects dataset preparation, molecular encoding, model configuration, training and candidate ranking for research applications such as repurposing and virtual screening.

For drug-target tasks, inputs are compound SMILES, protein amino acid sequences and measured binding values or binary labels. Other modules accept compound-only, protein-only or paired inputs with corresponding labels. Users can select fingerprint, descriptor, sequence or graph-based encodings, prepare training, validation and test partitions, and initialize a model or load a pretrained checkpoint. Outputs include predictions, training metrics, evaluation figures and ranked repurposing results. Dataset loaders cover sources including BindingDB, DAVIS, KIBA and the Broad Repurposing Hub; these are upstream data resources, not databases maintained by DeepPurpose.

The README documents regression and binary classification, cold-drug and cold-target evaluation settings, label-scale conversion and pretrained models associated with named datasets. It also cautions that the educational pretrained repurposing example has limited coverage and may not generalize to unseen proteins. BindingDB and DAVIS DTI checkpoints use log-scale labels, so output interpretation requires attention to conversion settings. The project calls for expert inspection before wet-lab validation and warns against directly using suggested drugs. These documented workflows support evaluation planning, not evidence of experimentally confirmed activity.

Key Features

  • Task modules for drug-target, drug-drug and protein-protein interaction prediction, compound property prediction and protein function prediction.
  • Selectable compound encodings include Morgan fingerprints, rdkit_2d_normalized descriptors, SMILES-based CNNs, transformers, MPNN and DGL graph models; protein encodings include sequence descriptors, CNNs and transformers.
  • Dataset loaders and text-file readers support benchmark binding data, bioassay labels, custom interaction pairs, repurposing libraries and virtual-screening inputs.
  • Data preparation supports cold-drug and cold-target splits, while training provides early stopping and classification or regression metrics.
  • Pretrained checkpoint loading, training from scratch and documented fine-tuning workflows support repurposing and virtual screening with ranked outputs.
  • Label conversion supports movement between logarithmic and original binding-value scales, including interpretation of pretrained DTI predictions.

Use Cases

  • Intended evaluation: compare compound and protein encodings for binding-affinity prediction using a documented benchmark and cold-drug or cold-target splits.
  • Intended evaluation: train a compound-property model on labeled screening data when target protein sequences are unavailable.
  • Intended evaluation: rank a supplied compound library against a target sequence for expert review and selection of candidates for experimental validation.
  • Intended evaluation: explore drug-drug or protein-protein interaction classification using labeled pairs in the documented input formats.

How to Use

  1. Start with the repository README to choose the task module and inspect its example. Match your problem to interaction regression, binary classification or a compound/protein prediction task.
  2. Follow the README’s installation instructions for local package installation or a source-based environment. Treat these as supplied setup guidance, not confirmation of compatibility with your environment.
  3. Consult the data-format descriptions and notebooks in the DEMO folder. Prepare SMILES, protein sequences and labels as required; check binding-value units and whether logarithmic conversion is appropriate.
  4. Select documented encodings and configure data partitions. For an intended generalization evaluation, consider cold-drug or cold-target splits rather than relying only on a random split. Train a model or consult the pretrained-model tutorial before choosing a checkpoint.
  5. Inspect task-appropriate metrics and prediction scales before running repurposing or virtual screening. Review ranked results with domain experts before experimental follow-up. The linked documentation site is described as under active development.

Related resources

AIDDISON Explorer is a hosted drug-discovery platform that generates and ranks molecular candidates against target profiles, design constraints, predicted properties and synthetic feasibility.

Molecular Generation · Drug Discovery

AutoDock Vina

Open Source

AutoDock Vina is an open-source molecular docking and virtual screening program with Vina and AutoDock4.2 scoring, multi-ligand docking, macrocycle support and Python 3 bindings.

Open sourcePython

Drug Discovery

Boltz

Model

Boltz is a biomolecular interaction model family. Boltz-2 provides workflows for complex structure and binding affinity prediction.

Open sourcePython

Drug Discovery

ChEMBL MCP (cyanheads) connects MCP clients to ChEMBL compound, target, bioactivity and drug records, with structure searches and optional DuckDB analysis of larger activity sets.

Open sourceTypeScript

Drug Discovery · Scientific Data

Official Python client maintained by the ChEMBL group for querying ChEMBL data and cheminformatics services, with QuerySet-style filters, lazy retrieval and local result caching.

Open sourcePython

Drug Discovery · Scientific Data

Chemistry42

Platforms

Chemistry42 is a small-molecule discovery platform combining generative design, retrosynthesis, ADMET and selectivity prediction, and physics-based prioritization for hit identification and lead optimization.

Molecular Generation · Drug Discovery

Related guides

Agents

AI Agents for Chemistry: From Chatbots to Autonomous Research

A practical guide to eight chemistry and biomedical research agents: how tools, memory, planning and multi-agent roles work, which projects fit different tasks, and how to evaluate bounded autonomy with scientific oversight.

Workflows

AI for Drug Discovery: Agents, MCP Servers and Platforms

Choose tools by workflow stage: target evidence, protein structures, known binding sites, molecular generation, property analysis and screening. Compare research agents, MCP integrations, hosted APIs and platforms, with a proposed end-to-end example and practical evaluation gates.

Workflows

Protein & Structural Biology AI Workflow

Build an evidence-led protein workflow from UniProt identity and sequences to RCSB structure metadata, observed ligand contacts, and proposed downstream candidate evaluation—with explicit checks for species, isoforms, chains, missing data, and experimental validation.