Overview
GeoDiff is the official implementation of a geometric diffusion model for molecular conformation generation described in an ICLR 2022 paper. Its documented workflow centers on learning from GEOM molecular datasets and generating conformations for test-set molecules. It is a research model implementation with training, sampling, and evaluation scripts, rather than an upstream molecular database or a general-purpose chemistry toolkit.
The repository links to both the original GEOM dataset and separately supplied preprocessed datasets. Dataset locations and training settings are configured through YAML files. Training produces checkpoints, a configuration file, and logs in a selected output directory. For generation, the project provides qm9_default and drugs_default checkpoints and supports selecting all or part of a test set. Generated conformations are subsequently consumed by evaluation scripts, with examples referencing sample_all.pkl files. The documented evaluations include GEOM conformation coverage and matching scores, plus a property-prediction workflow using a separate small QM9 split.
Branch and dependency choices are important limitations. The README directs users of the supplied pretrained checkpoints to the pretrain branch because changes on main introduced compatibility problems with those checkpoints. It also reports that both the codebase and processed PyD data objects are incompatible with more recent torch-geometric versions. The supplied material therefore supports a specific research workflow, not a claim of compatibility with current environments or validated performance on arbitrary molecules.
Key Features
- Configuration-driven training on GEOM datasets, including QM9, drug-like molecules, and a documented reduced-timestep ablation setting.
- Supplied `qm9_default` and `drugs_default` pretrained checkpoints, with a designated `pretrain` branch for their use.
- Conformation generation for complete test sets or index-selected subsets, with sampling settings exposed in `test.py`.
- Conformation evaluation through `eval_covmat.py`, reporting `COV` and `MAT` scores on GEOM datasets.
- A separate QM9 property-prediction evaluation workflow using generated conformations and `eval_prop.py`.
- Guidance for displaying and aligning molecular conformations in PyMol.
Use Cases
- Intended evaluation: generate conformations for a bounded subset of the supplied GEOM test data and assess them with the documented `COV` and `MAT` workflow.
- Intended evaluation: compare configuration-controlled training settings, including the reduced-timestep drug-model ablation.
- Intended evaluation: investigate how generated conformations support the repository’s property-prediction benchmark on its separate QM9 split.
- Intended evaluation: visually inspect generated conformations using the supplied PyMol display and alignment guidance.
How to Use
- Read the official README, focusing on its environment instructions and known torch-geometric incompatibilities before choosing a local setup.
- Obtain the preprocessed GEOM data from the linked download folder. Place it at the locations specified by the configuration files’
datasetsettings; retain the distinction from the original GEOM dataset. - Choose between training and checkpoint-based generation. For supplied pretrained checkpoints, use the pretrain branch and preserve the documented checkpoint/configuration directory layout.
- For training, inspect the repository’s YAML settings and
train.pyexamples. Verify configuration paths before proceeding: the README refers to both./configs/*.ymland./config/examples. - For an initial evaluation, select a small test-set range in the documented
test.pyworkflow, then inspect its generated output witheval_covmat.py. Treat the separate QM9 property split as a distinct evaluation, and verify that example against the code before use.