Skip to content

TorchDrug

TorchDrug Team

TorchDrug is a PyTorch-based toolkit for graph and molecular machine learning, with SMILES input, graph operations, property-prediction building blocks and multi-CPU/GPU execution.

Catalog updated ·

Overview

TorchDrug is a machine learning toolbox for developing graph-based drug-discovery workflows in PyTorch. It supplies data representations, graph operations, datasets, models and task components rather than a single pretrained predictor. Its documented role is to support research prototyping and the assembly of training and inference workflows from reusable building blocks.

Inputs illustrated in the README include graph edge lists and molecular SMILES strings. Graphs can be moved to a GPU and reduced to node-induced subgraphs. Molecular objects expose atom types and node features and can produce scaffold representations. Users can also attach custom node, edge or graph attributes that are handled during indexing, making these objects useful for experiments requiring additional annotations.

The property-prediction example loads Tox21, creates training, validation and test subsets, and connects a GIN model to a PropertyPrediction task. The Engine component supports CPU, GPU and distributed execution, while an optional Weights & Biases logger provides experiment tracking. Outputs illustrated by the source include molecular features, scaffolds and subgraphs; prediction behavior depends on the task and model assembled by the user.

The supplied installation guidance specifies Python 3.7–3.10 and PyTorch requirements, with additional setup considerations for Windows and Apple silicon. It explicitly states that mps devices are unsupported. Although the README links to benchmarks, the source excerpts contain no benchmark results or evidence of predictive accuracy for a particular scientific application.

Key Features

  • PyTorch-backed graph operations with GPU acceleration and automatic differentiation, including node-induced subgraph extraction.
  • Molecule construction from SMILES with default atom and bond features, atom-type access and scaffold conversion.
  • Custom node, edge and graph attributes that are automatically handled during indexing operations.
  • Dataset and modeling components illustrated by Tox21, GIN and the PropertyPrediction task.
  • An Engine interface for training and inference using single or multiple CPUs and GPUs, including distributed configurations.
  • Optional experiment logging through Weights & Biases.

Use Cases

  • Suggested evaluation: build a Tox21 property-prediction baseline using the documented GIN and PropertyPrediction components, then assess it on a held-out subset.
  • Suggested evaluation: convert representative SMILES strings into molecular objects and inspect atom features and scaffolds for a molecular representation workflow.
  • Suggested evaluation: attach custom graph annotations and check that subgraph extraction preserves the expected attributes.
  • Suggested evaluation: compare available CPU and GPU execution configurations for a chosen training workflow; no speedup is established by the source excerpts.

How to Use

  1. Read the repository README and select its Conda, pip or source installation route. Check the stated Python and PyTorch requirements; consult the Windows or Apple silicon notes if applicable.
  2. Follow the documentation to choose a graph or molecular representation. Start with a small edge list or representative SMILES strings and inspect the resulting features before proceeding.
  3. Use the tutorials alongside the README examples to explore subgraphs, scaffolds and custom attributes. Check indexing behavior on your own annotations.
  4. Assemble a property-prediction workflow using the documented Tox21, GIN and PropertyPrediction example as a starting point. Define training, validation and test subsets appropriate to your intended evaluation.
  5. Configure Engine for the available CPU or GPU resources and optionally enable Weights & Biases logging. Consult the linked benchmarks for further context, then evaluate predictive quality and resource use independently.

Related resources

DrugAgent is a research multi-agent framework that combines LLM planning with domain-guided code generation to build and evaluate machine-learning workflows for drug discovery.

Molecular Property Prediction · Drug Discovery

Therapeutics Data Commons provides therapeutic machine learning datasets, Python loaders, data splits, evaluation metrics and benchmarks for prediction and molecule-generation research.

Open sourcePython

Molecular Property Prediction · Drug Discovery

AIDDISON Explorer is a hosted drug-discovery platform that generates and ranks molecular candidates against target profiles, design constraints, predicted properties and synthetic feasibility.

Molecular Generation · Drug Discovery

AutoDock Vina

Open Source

AutoDock Vina is an open-source molecular docking and virtual screening program with Vina and AutoDock4.2 scoring, multi-ligand docking, macrocycle support and Python 3 bindings.

Open sourcePython

Drug Discovery

Boltz

Model

Boltz is a biomolecular interaction model family. Boltz-2 provides workflows for complex structure and binding affinity prediction.

Open sourcePython

Drug Discovery

ChemBERTa provides BERT-like models for chemical SMILES, with RoBERTa masked-language modelling checkpoints and notebooks for pre-training, fine-tuning and molecular property prediction research.

Open sourcePython

Molecular Property Prediction

Related guides

Workflows

AI for Drug Discovery: Agents, MCP Servers and Platforms

Choose tools by workflow stage: target evidence, protein structures, known binding sites, molecular generation, property analysis and screening. Compare research agents, MCP integrations, hosted APIs and platforms, with a proposed end-to-end example and practical evaluation gates.