Skip to content

Matbench Discovery

Janosh Riebesell

Matbench Discovery benchmarks machine-learning models for crystal stability and atomistic simulation tasks, using an interactive leaderboard to compare accuracy, robustness, and computational cost.

Catalog updated ·

Overview

Matbench Discovery is a benchmark and interactive leaderboard for assessing machine-learning models in inorganic crystal discovery and atomistic simulation. Its documented scope extends beyond stability prediction to geometry optimization, phonons and thermal conductivity, molecular dynamics, and diatomic potential-energy curves. It supports model comparison rather than serving as a single predictive model or a general-purpose materials database.

The leaderboard includes graph neural network interatomic potentials, graph neural network one-shot predictors, iterative Bayesian optimizers, and random forests using shallow-learning structure fingerprints. In a research workflow, it can inform which approaches merit further evaluation for static or finite-temperature simulations. The benchmark evaluates predicted material properties, including thermodynamic stability, thermal conductivity, and atomic positions, against task-specific reference data. Its outputs include model rankings and comparisons of accuracy, robustness, and computational cost; the source excerpts do not specify complete input schemas or submission artifact formats.

A central interpretation constraint is that crystal stability is evaluated using a convex hull constructed from DFT reference energies, not model-predicted energies. Users should account for that choice when interpreting results or comparing them with other benchmarks. The project also cautions that leaderboard position does not fully characterize a model’s ability to advance materials research and does not imply Materials Project endorsement. A contributing guide provides the documented route for submitting new models, while the repository includes Python package configuration.

Key Features

  • Interactive leaderboard ranking models across crystal stability, geometry optimization, phonons and thermal conductivity, molecular dynamics, and diatomic potential-energy curves.
  • Coverage of multiple model families, including GNN interatomic potentials, GNN one-shot predictors, iterative Bayesian optimizers, and fingerprint-based random forests.
  • Task-specific comparisons intended to expose accuracy, robustness, and computational-cost trade-offs.
  • Stability evaluation against a convex hull constructed from DFT reference energies rather than model predictions.
  • A model-contribution workflow with a linked guide and branch-specific pull-request instructions.

Use Cases

  • Suggested evaluation: shortlist models for crystal stability screening, then assess their suitability for the intended discovery workflow.
  • Suggested evaluation: compare candidate interatomic potentials for geometry optimization or finite-temperature simulation using the relevant leaderboard tasks.
  • Suggested evaluation: examine phonon and thermal-conductivity results before planning independent validation on a target materials system.
  • Suggested evaluation: prepare a new model submission using the contribution guide to compare it within the benchmark’s reference-data framework.

How to Use

  1. Open the interactive leaderboard and identify the task relevant to your work, such as crystal stability, geometry optimization, or thermal conductivity.
  2. Consult the stability benchmark explanation before interpreting discovery rankings. Note that the evaluation hull uses DFT reference energies rather than model predictions.
  3. Compare the relevant model families and available accuracy, robustness, and computational-cost information. Treat this as a shortlist for further evaluation, not proof of suitability for your materials or simulation conditions.
  4. If preparing local benchmark work, inspect the repository configuration for declared dependencies and Python requirements. Confirm task-specific inputs and procedures in the project documentation; the source excerpts do not provide an execution recipe.
  5. For a new model submission, follow the contributing guide, including its branch-specific instructions. Use GitHub discussions for support questions.

Related resources

ALIGNN

Model

ALIGNN provides atomistic graph neural networks for materials property prediction, with training workflows, pretrained predictors and ALIGNN-FF force fields for structural optimization.

Open sourcePython

Materials Discovery

Allegro

Model

Allegro implements an E(3)-equivariant interatomic potential as a NequIP extension, with documented GPU acceleration options and a separate plugin for LAMMPS simulations.

Open sourcePython

Computational Chemistry · Materials Discovery

An ASE routing Skill in the computational-chemistry-agent-skills collection that separates workflow preparation from calculator configuration and delegates execution elsewhere.

Computational Chemistry · Materials Discovery

CGCNN

Model

CGCNN implements crystal graph convolutional neural networks for learning material properties from crystal structures, with custom-data training and prediction using pre-trained models.

Open sourcePython

Materials Discovery

ChatMOF

Agent

ChatMOF is a code-available research system that uses language-model-guided tools to retrieve MOF data, predict properties and generate structures from natural-language requests.

Open sourcePython

Materials Discovery

ChemAgent (AI4Chem) is a research framework for chemistry and materials tool use, linked to the CheMatAgent paper on tree-search planning, tool execution and ChemToolBench-based training.

Computational Chemistry · Materials Discovery

Related guides

Agents

AI Agents for Chemistry: From Chatbots to Autonomous Research

A practical guide to eight chemistry and biomedical research agents: how tools, memory, planning and multi-agent roles work, which projects fit different tasks, and how to evaluate bounded autonomy with scientific oversight.