Overview
DimeNet is a research repository containing reference implementations of the DimeNet and DimeNet++ neural-network models for molecular graphs. It accompanies the papers Directional Message Passing for Molecular Graphs and Fast and Uncertainty-Aware Directional Message Passing for Non-Equilibrium Molecules. Its workflow role is model training and molecular-property prediction, rather than molecular database access or a hosted prediction service. The README recommends DimeNet++ over the original model.
The models use molecular graph information with distances and angles in their directional message-passing calculations. The documented QM9 targets include alpha, R2, U0, U, H, G, Cv, Mu, HOMO, LUMO and ZPVE. The source excerpts do not specify a complete input schema or prediction-file format. Training is organized through train.ipynb, while predict.ipynb generates test-set predictions using a trained model. Two sets of pretrained models are provided in pretrained, and train_seml.py supports cluster training with Sacred and SEML.
This repository uses TensorFlow 2 and documents training and initialization differences from the original TensorFlow 1 implementation. Interpretation of results requires attention to the listed issues: an off-by-one molecule-filtering error in qm9_eV.npz, distance and basis-layer discrepancies, an embedding-layer discrepancy, and a checkpointing problem affecting older TensorFlow AddOns versions. The README also identifies an MD17 benchmark-comparison mismatch. For energy and force prediction, the authors instead recommend GemNet; this entry should therefore be treated as a documented research implementation, not an independently validated scientific workflow.
Key Features
- Includes reference implementations of both DimeNet and DimeNet++ for directional message passing on molecular graphs.
- Provides `train.ipynb` for model training and `predict.ipynb` for test-set prediction with a trained model.
- Supplies two sets of pretrained models in the `pretrained` folder for experimentation.
- Includes `train_seml.py` for cluster training using Sacred and SEML.
- Documents target-specific output-layer initialization in the TensorFlow 2 implementation.
- Lists known dataset, geometry, basis-layer, embedding-layer and checkpointing issues relevant to interpreting or reproducing results.
Use Cases
- Intended evaluation: use the training and prediction notebooks to assess a QM9 molecular-property prediction workflow, accounting for the documented molecule-filtering error.
- Intended evaluation: compare DimeNet and DimeNet++ within a controlled experimental setup without assuming that published results transfer to a new dataset.
- Intended evaluation: examine the supplied pretrained models before committing resources to retraining.
- Intended evaluation: assess `train_seml.py` as a starting point for cluster-based molecular-model experiments using Sacred and SEML.
How to Use
- Read the official README to choose between DimeNet and DimeNet++. It recommends DimeNet++ and points energy-and-force users toward GemNet.
- Review the dependencies in setup.py, including TensorFlow, TensorFlow AddOns, NumPy, SciPy and SymPy. Treat these as declared requirements, not evidence of compatibility with your environment.
- Inspect
train.ipynb,predict.ipynband thepretrainedfolder in the repository. Check the expected data representation and target selection before attempting training or prediction. - Review the README's known issues and the QM9 filtering discussion. For an intended evaluation, record dataset provenance, initialization choices and any corrections separately from upstream results.
- If cluster training is needed, examine
train_seml.pyand the linked SEML project. Review the code licence before use or redistribution.