Overview
DiffDock is a research implementation for predicting the three-dimensional structure of a small molecule bound to a protein. It supports the pose-generation stage of a molecular docking workflow rather than directly estimating binding affinity. The project README states that the repository runs DiffDock-L by default and directs users seeking the original DiffDock model to the repository history.
Inputs can combine a protein structure in .pdb format with a ligand represented as a SMILES string or an RDKit-readable file, such as .sdf or .mol2. Alternatively, users can provide a protein sequence, which the workflow folds with ESMFold. Individual complexes can be submitted directly, while a CSV describes multiple protein–ligand pairs for batch inference. Outputs include predicted binding structures and confidence scores. The repository also documents a local graphical interface for single complexes and links to a hosted Hugging Face Spaces interface.
Confidence scores describe the model’s confidence in a predicted pose, not binding affinity, and the README cautions against straightforward comparisons across complexes or protein conformations. Its interpretation guidance assumes inputs resembling the training setting: drug-like ligands and medium-sized proteins in near-bound conformations. The model was designed for small-molecule docking to proteins, not interactions between larger biomolecules. For downstream affinity estimation, the authors suggest relaxing predicted structures before applying separate scoring or free-energy methods. Local inference with supplied .pdb structures supports CPU execution, although the README recommends a GPU.
Key Features
- Predicts protein–small-molecule complex structures and provides pose confidence scores; the current repository defaults to DiffDock-L.
- Accepts protein `.pdb` files or protein sequences folded through ESMFold.
- Accepts ligand SMILES strings and RDKit-readable molecular files, including `.sdf` and `.mol2`.
- Supports single-complex inference and batch inference from CSV records containing protein and ligand inputs.
- Provides a local single-complex graphical interface, a hosted Hugging Face Spaces interface, and documented Conda and Docker setup routes.
- Documents benchmark evaluation workflows for PDBBind, DockGen, and PoseBusters, including preparation of cached ESM2 embeddings.
Use Cases
- Intended evaluation: generate candidate binding poses for a drug-like ligand and a protein structure, then inspect the poses and their confidence scores.
- Intended evaluation: process a collection of protein–ligand pairs through CSV-based inference to assess suitability for a research docking workflow.
- Intended evaluation: explore sequence-based docking through ESMFold-generated protein structures while accounting for the README’s cautions about unbound conformations.
- Intended evaluation: use predicted structures as starting inputs for separate affinity-scoring or free-energy workflows, with relaxation as recommended by the authors.
How to Use
- Read the official repository documentation, especially the inference and FAQ sections. Decide whether you need the default DiffDock-L model or the original model described through repository history.
- For an initial single-complex trial, use the linked Hugging Face Spaces interface. For local work, follow the README’s Conda or Docker setup instructions; it also documents a local graphical interface.
- Prepare a protein
.pdbfile or a sequence for ESMFold, plus a ligand SMILES string or RDKit-readable file. For multiple complexes, follow the documented CSV fields and consultdata/protein_ligand_example.csv. - Follow the repository’s inference example to select the configuration, inputs, and output directory. Account for initial distribution-table caching; the README recommends GPU execution where available.
- Inspect predicted structures and confidence scores using the FAQ’s qualifications. For a proposed downstream affinity workflow, evaluate relaxation and separate scoring methods rather than treating confidence as affinity. Consult the repository’s dataset and replication sections for benchmark-oriented evaluation.