Skip to content

Therapeutics Data Commons

PyTDC Team

Therapeutics Data Commons provides therapeutic machine learning datasets, Python loaders, data splits, evaluation metrics and benchmarks for prediction and molecule-generation research.

Catalog updated ·

Overview

Therapeutics Data Commons (TDC) organizes datasets and evaluation workflows for machine learning across therapeutic discovery and development. Its documented scope includes target discovery, activity screening, efficacy, safety and manufacturing, with biomedical products such as small molecules, antibodies and vaccines. The resource combines dataset access with supporting tools for model development and comparison; it is not a single predictive model.

TDC arranges its resources into problems, learning tasks and datasets. The problem categories distinguish prediction for individual biomedical entities, prediction involving multiple entities, and generation of new entities. Researchers select a task and dataset through Python APIs, retrieve data in supported formats, and obtain training, validation and test partitions. The README demonstrates loading the ADME dataset HIA_Hou, returning a dataframe and constructing a scaffold split. Evaluation utilities take reference labels and model predictions to calculate a selected metric, while molecule-generation oracles return scores for supplied molecules.

Within a research workflow, TDC can provide the data preparation and evaluation layer around a separately developed model. Processing functions cover label transformation, balancing, negative sampling and graph preparation, while themed benchmark groups support comparisons across related datasets, including ADMET tasks. The README also links tutorials for single-cell resources, PrimeKG, external APIs and a model hub.

The project README describes the package as a beta release. Dataset retrieval also depends on hosting availability: many datasets reside on Harvard Dataverse, and its maintenance can interrupt access. The repository's MIT code licence does not establish dataset usage rights; those must be checked individually. Documented benchmark facilities should not be interpreted as evidence that a particular model is suitable for clinical use.

Key Features

  • Hierarchical Python data loaders organize access by prediction or generation problem, therapeutic learning task and dataset.
  • Dataset partitioning supports configurable split methods, random seeds and fractions, including the documented scaffold-split example.
  • Evaluation utilities calculate task metrics such as ROC-AUC from reference labels and model predictions.
  • Data-processing functions include label transformation, balancing, negative sampling and preparation of PyG/DGL graphs.
  • Molecule-generation oracles support goal-oriented and distribution-learning workflows, including the documented GSK3B oracle.
  • Themed benchmark groups provide dataset access, evaluation of predictions across multiple runs and a workflow for leaderboard submissions.

Use Cases

  • Suggested evaluation: compare molecular-property models on an ADME dataset using a scaffold split to examine generalization to unseen compounds.
  • Suggested evaluation: assess a model across the ADMET_Group benchmarks using the documented partitions and evaluation workflow.
  • Suggested evaluation: use TDC molecule-generation oracles to score candidate molecules during a generation experiment, without treating oracle scores as experimental validation.
  • Suggested evaluation: prepare therapeutic datasets for graph-learning experiments using the documented PyG/DGL processing functions.

How to Use

  1. Start with the task and dataset overview to select a therapeutic question and inspect the relevant dataset description, original publication and dataset-specific licence.

  2. Follow the installation guidance in the repository README. It identifies PyTDC as the package and describes the release as beta; consult the documentation for API guidance.

  3. Use the documented loader pattern to retrieve your chosen dataset. Inspect the returned data before modeling; the README illustrates HIA_Hou retrieval as a dataframe rather than providing a trained predictor.

  4. Choose an appropriate partition using the data split guidance. Record the split method, seed and fractions so comparisons use the same evaluation setup.

  5. Train your own model and calculate the relevant metric following the evaluation guidance. For grouped comparisons or submissions, follow the benchmark overview.

  6. If retrieval fails, check Harvard Dataverse availability; the README identifies maintenance there as a possible interruption.

Related resources

DrugAgent is a research multi-agent framework that combines LLM planning with domain-guided code generation to build and evaluate machine-learning workflows for drug discovery.

Molecular Property Prediction · Drug Discovery

TorchDrug

Open Source

TorchDrug is a PyTorch-based toolkit for graph and molecular machine learning, with SMILES input, graph operations, property-prediction building blocks and multi-CPU/GPU execution.

Open sourcePython

Molecular Property Prediction · Drug Discovery

AIDDISON Explorer is a hosted drug-discovery platform that generates and ranks molecular candidates against target profiles, design constraints, predicted properties and synthetic feasibility.

Molecular Generation · Drug Discovery

AutoDock Vina

Open Source

AutoDock Vina is an open-source molecular docking and virtual screening program with Vina and AutoDock4.2 scoring, multi-ligand docking, macrocycle support and Python 3 bindings.

Open sourcePython

Drug Discovery

Boltz

Model

Boltz is a biomolecular interaction model family. Boltz-2 provides workflows for complex structure and binding affinity prediction.

Open sourcePython

Drug Discovery

ChemBERTa provides BERT-like models for chemical SMILES, with RoBERTa masked-language modelling checkpoints and notebooks for pre-training, fine-tuning and molecular property prediction research.

Open sourcePython

Molecular Property Prediction

Related guides

Overview

How to Build a Chemistry AI Agent Stack

Design a chemistry AI agent stack by separating language models, orchestration, Skills, MCP interfaces, APIs, libraries and data. Follow proposed molecular and materials workflows, define input/output contracts, and plan permissions, evaluation and reproducible deployment.

Licensing

Reading Code, Weight and Data Licences Separately

Build a component-by-component licensing record for chemistry AI workflows, separating software, model weights, datasets and hosted services while keeping missing evidence and unresolved conditions visible.