Overview
Matbench Discovery is a benchmark and interactive leaderboard for assessing machine-learning models in inorganic crystal discovery and atomistic simulation. Its documented scope extends beyond stability prediction to geometry optimization, phonons and thermal conductivity, molecular dynamics, and diatomic potential-energy curves. It supports model comparison rather than serving as a single predictive model or a general-purpose materials database.
The leaderboard includes graph neural network interatomic potentials, graph neural network one-shot predictors, iterative Bayesian optimizers, and random forests using shallow-learning structure fingerprints. In a research workflow, it can inform which approaches merit further evaluation for static or finite-temperature simulations. The benchmark evaluates predicted material properties, including thermodynamic stability, thermal conductivity, and atomic positions, against task-specific reference data. Its outputs include model rankings and comparisons of accuracy, robustness, and computational cost; the source excerpts do not specify complete input schemas or submission artifact formats.
A central interpretation constraint is that crystal stability is evaluated using a convex hull constructed from DFT reference energies, not model-predicted energies. Users should account for that choice when interpreting results or comparing them with other benchmarks. The project also cautions that leaderboard position does not fully characterize a model’s ability to advance materials research and does not imply Materials Project endorsement. A contributing guide provides the documented route for submitting new models, while the repository includes Python package configuration.
Key Features
- Interactive leaderboard ranking models across crystal stability, geometry optimization, phonons and thermal conductivity, molecular dynamics, and diatomic potential-energy curves.
- Coverage of multiple model families, including GNN interatomic potentials, GNN one-shot predictors, iterative Bayesian optimizers, and fingerprint-based random forests.
- Task-specific comparisons intended to expose accuracy, robustness, and computational-cost trade-offs.
- Stability evaluation against a convex hull constructed from DFT reference energies rather than model predictions.
- A model-contribution workflow with a linked guide and branch-specific pull-request instructions.
Use Cases
- Suggested evaluation: shortlist models for crystal stability screening, then assess their suitability for the intended discovery workflow.
- Suggested evaluation: compare candidate interatomic potentials for geometry optimization or finite-temperature simulation using the relevant leaderboard tasks.
- Suggested evaluation: examine phonon and thermal-conductivity results before planning independent validation on a target materials system.
- Suggested evaluation: prepare a new model submission using the contribution guide to compare it within the benchmark’s reference-data framework.
How to Use
- Open the interactive leaderboard and identify the task relevant to your work, such as crystal stability, geometry optimization, or thermal conductivity.
- Consult the stability benchmark explanation before interpreting discovery rankings. Note that the evaluation hull uses DFT reference energies rather than model predictions.
- Compare the relevant model families and available accuracy, robustness, and computational-cost information. Treat this as a shortlist for further evaluation, not proof of suitability for your materials or simulation conditions.
- If preparing local benchmark work, inspect the repository configuration for declared dependencies and Python requirements. Confirm task-specific inputs and procedures in the project documentation; the source excerpts do not provide an execution recipe.
- For a new model submission, follow the contributing guide, including its branch-specific instructions. Use GitHub discussions for support questions.