Overview
DeepPurpose is a molecular modeling toolkit centered on drug-target interaction prediction, with additional modules for compound properties, drug-drug interactions, protein-protein interactions and protein function. It provides a configurable training and inference framework rather than a single predictive model. Its documented workflow connects dataset preparation, molecular encoding, model configuration, training and candidate ranking for research applications such as repurposing and virtual screening.
For drug-target tasks, inputs are compound SMILES, protein amino acid sequences and measured binding values or binary labels. Other modules accept compound-only, protein-only or paired inputs with corresponding labels. Users can select fingerprint, descriptor, sequence or graph-based encodings, prepare training, validation and test partitions, and initialize a model or load a pretrained checkpoint. Outputs include predictions, training metrics, evaluation figures and ranked repurposing results. Dataset loaders cover sources including BindingDB, DAVIS, KIBA and the Broad Repurposing Hub; these are upstream data resources, not databases maintained by DeepPurpose.
The README documents regression and binary classification, cold-drug and cold-target evaluation settings, label-scale conversion and pretrained models associated with named datasets. It also cautions that the educational pretrained repurposing example has limited coverage and may not generalize to unseen proteins. BindingDB and DAVIS DTI checkpoints use log-scale labels, so output interpretation requires attention to conversion settings. The project calls for expert inspection before wet-lab validation and warns against directly using suggested drugs. These documented workflows support evaluation planning, not evidence of experimentally confirmed activity.
Key Features
- Task modules for drug-target, drug-drug and protein-protein interaction prediction, compound property prediction and protein function prediction.
- Selectable compound encodings include Morgan fingerprints, rdkit_2d_normalized descriptors, SMILES-based CNNs, transformers, MPNN and DGL graph models; protein encodings include sequence descriptors, CNNs and transformers.
- Dataset loaders and text-file readers support benchmark binding data, bioassay labels, custom interaction pairs, repurposing libraries and virtual-screening inputs.
- Data preparation supports cold-drug and cold-target splits, while training provides early stopping and classification or regression metrics.
- Pretrained checkpoint loading, training from scratch and documented fine-tuning workflows support repurposing and virtual screening with ranked outputs.
- Label conversion supports movement between logarithmic and original binding-value scales, including interpretation of pretrained DTI predictions.
Use Cases
- Intended evaluation: compare compound and protein encodings for binding-affinity prediction using a documented benchmark and cold-drug or cold-target splits.
- Intended evaluation: train a compound-property model on labeled screening data when target protein sequences are unavailable.
- Intended evaluation: rank a supplied compound library against a target sequence for expert review and selection of candidates for experimental validation.
- Intended evaluation: explore drug-drug or protein-protein interaction classification using labeled pairs in the documented input formats.
How to Use
- Start with the repository README to choose the task module and inspect its example. Match your problem to interaction regression, binary classification or a compound/protein prediction task.
- Follow the README’s installation instructions for local package installation or a source-based environment. Treat these as supplied setup guidance, not confirmation of compatibility with your environment.
- Consult the data-format descriptions and notebooks in the DEMO folder. Prepare SMILES, protein sequences and labels as required; check binding-value units and whether logarithmic conversion is appropriate.
- Select documented encodings and configure data partitions. For an intended generalization evaluation, consider cold-drug or cold-target splits rather than relying only on a random split. Train a model or consult the pretrained-model tutorial before choosing a checkpoint.
- Inspect task-appropriate metrics and prediction scales before running repurposing or virtual screening. Review ranked results with domain experts before experimental follow-up. The linked documentation site is described as under active development.