Overview
EquiBind is a geometric deep learning model for protein–ligand binding structure prediction. Its SE(3)-equivariant architecture predicts both the receptor binding location, described as blind docking, and the ligand’s bound pose and orientation in a direct-shot workflow. The repository provides model weights, inference configurations and training instructions, making it a resource for generating structural predictions rather than a database of experimentally established binding poses.
For standard inference, each protein–ligand pair occupies a separate input directory. Receptors are supplied as .pdb files with protein in their filenames; ligands may use .mol2, .sdf, .pdbqt or .pdb, with ligand in their filenames and all hydrogens included. Predictions are written as .sdf files and tensors. A separate workflow processes multiple ligands from one .sdf file against a single receptor, producing predicted conformers alongside lists of successful and handled failed cases.
The documented setup offers separate CUDA GPU and CPU-only environment files. Researchers can also train a model, save weights in runs, and evaluate those weights through the reproduction configuration. These paths support evaluation of the model within a structure-prediction workflow, but the source excerpts do not establish accuracy for a particular target or dataset.
Dataset access is a practical constraint: the previously hosted preprocessed PDBBind data is no longer available because redistribution is not permitted under its licence. Users must obtain it from PDBBind directly for the documented dataset workflow. The README also recommends considering DiffDock as a newer approach; this entry does not independently assess that comparison.
Key Features
- Direct-shot prediction of both receptor binding location and ligand bound pose using an SE(3)-equivariant geometric model.
- Provided-weight inference for protein–ligand pairs, accepting four documented ligand file formats and `.pdb` receptors.
- Prediction export as `.sdf` structures and tensors, with output locations controlled by the inference configuration.
- Multi-ligand inference against one receptor, with conformer output and separate success and handled-failure records.
- Separate documented CUDA GPU and CPU-only environment configurations.
- Training and test-set evaluation workflows, including reproduction configurations for EquiBind-U and EquiBind with ligand point-cloud fitting corrections.
Use Cases
- Intended evaluation: generate candidate binding locations and ligand poses for prepared protein–ligand pairs, then assess them against appropriate structural references.
- Intended evaluation: process a multi-ligand `.sdf` file against one receptor and inspect both predicted conformers and recorded inference failures.
- Intended evaluation: train on appropriately obtained PDBBind data and compare resulting predictions with those from the supplied weights using the documented test workflow.
How to Use
- Read the official README and consult the paper for methodological context. Choose between provided-weight inference and a training or reproduction workflow.
- Obtain the repository and follow its Anaconda setup instructions. Select
environment.ymlfor the documented CUDA path orenvironment_cpuonly.ymlfor CPU-only setup; verify suitability for your own environment. - Prepare one directory per complex. Use a receptor
.pdbfilename containingproteinand a supported ligand filename containingligand. Include all ligand hydrogens; the README notes use of reduce on training proteins. - Set
inference_pathand checkoutput_directoryinconfigs_clean/inference.yml. Follow the README’s inference procedure, then inspect the saved.sdfstructures and tensor output. For a multi-ligand input, inspect its success and failure records as well. - For training or reproduction, obtain data directly from PDBBind and place it at
data/PDBBind. Follow the documented configurations and assess predictions against suitable references before downstream use.