Define the reference problem first
An atomistic potential evaluation should begin with a scientific specification, not a repository ranking. This collection brings together four candidate projects for computational chemistry and materials discovery. Their workflows, checkpoints and simulation interfaces should not be treated as interchangeable.
Write a one-page specification covering elements, charge and spin where relevant, molecular or periodic structures, geometry range, and intended use. Distinguish single-point prediction, geometry optimisation and molecular dynamics: each needs its own acceptance criteria.
Record the reference calculation procedure, including energy units, coordinate units, force convention, boundary conditions and any energy corrections. Decide whether you need total energies, relative energies, forces, stress or additional properties. Keep this operating envelope attached to every comparison so that a result for one domain is not presented as evidence for another.
Inspect the four projects independently
The following capabilities are described by the project sources; they are not findings from ChemAI Atlas testing.
- MACE provides higher-order equivariant interatomic potentials, training and XYZ-based evaluation workflows, pretrained model inference and fine-tuning. Its ASE examples demonstrate calculator use. The sources distinguish materials-focused MACE-MP from organic-chemistry MACE-OFF models, making model-family selection part of the evaluation.
- NequIP is a framework for E(3)-equivariant interatomic potentials. It documents foundation potentials, compilation, fine-tuning, and ASE and LAMMPS interfaces. Allegro is a separate extension package, not a synonym for the core framework. Exact data schemas and prediction interfaces remain questions for the relevant guides.
- SchNetPack supports constructing and training atomistic networks using SchNet and PaiNN representations. Its examples cover quantum-chemical properties and energy surfaces with force and stress derivatives. It also describes molecular dynamics components and a LAMMPS interface. Select a model configuration rather than treating the toolbox as one pretrained potential.
- TorchANI supports developing, training and using ANI-style potentials in PyTorch, with optional C++ and CUDA extensions. Its README points to a separate TorchANI-Amber interface for full machine-learning and ML/MM simulations. Individual model coverage, input schemas and output structures are not established by the source excerpts.
For each candidate, ask: does the specific model address the target chemistry, or does the project merely provide infrastructure for training one?
Build a comparison worksheet
Create one row per selected release and model asset, not just per repository. Use the following fields to make missing evidence visible:
| Field | What to record |
|---|---|
| Identity | Release or commit, model identifier and asset provenance |
| Chemical coverage | Elements, systems, charge/spin conditions and documented exclusions |
| Reference data | Training source, level of theory and energy conventions |
| Inputs | Geometry schema, labels, units, cells and preprocessing |
| Outputs | Property names, shapes, units and derivative conventions |
| Training | Split files, seed, losses, settings and reference-energy treatment |
| Deployment | Calculator or simulation interface, compilation and dependencies |
| Evidence status | Source-described, proposed for evaluation, locally verified, or unknown |
Do not populate the “locally verified” status until your team has actually performed and recorded the relevant check. An undocumented condition should remain unknown rather than becoming an assumed default.
Use decision questions to narrow the shortlist: are required labels available? Is a matching pretrained asset documented? Can the intended output be obtained through the chosen interface? Would adapting a model change the scope or budget of the comparison?
Align inputs and outputs before scoring
The steps below are a proposed evaluation procedure, not a tested cross-project integration.
- Freeze a small reference set with stable configuration identifiers and preserve the original calculation records.
- Prepare each project's required input independently. Check element encoding, atom order, coordinates, cell information and property labels after conversion.
- Document output mappings to a common comparison table: configuration identifier, predicted and reference energy, and atom-indexed force components. Include stress only where it is required and supported by the selected setup.
- Check units, force signs, energy offsets and total-versus-per-atom reporting before calculating errors.
- Keep conversion failures, unsupported inputs and non-finite predictions in a separate failure log; do not silently remove them from the denominator.
MACE documents evaluation from an XYZ file and saved model to an output XYZ file. That does not establish an identical file contract for the other projects. Likewise, a documented simulation interface does not prove compatibility with your installed environment.
Worked planning example: neutral molecular conformers
Hypothetical example: a team wants relative conformer energies and forces for neutral molecules containing carbon, hydrogen, nitrogen and oxygen. This is a planning scenario, not a performance result.
The team would first specify a single reference electronic-structure procedure and assemble a labelled configuration inventory. It could separate conformer families across training, validation and test sets rather than rely only on random frame assignment. A separately labelled challenge set might contain more distorted geometries outside the intended operating envelope.
For MACE, the team would investigate an organic model asset or a training workflow and check its reference chemistry. For NequIP, it would establish whether a documented pretrained asset fits the specification or whether training is needed. For SchNetPack, it would select a representation and energy/force output configuration. For TorchANI, it would first confirm the chosen model's elements and molecular domain in its documentation.
The proposed output is a coverage worksheet, aligned prediction table and failure report. If an asset's energy reference differs from the team's reference method, the team should resolve that mismatch or evaluate it as a separate question—not declare it inferior from raw absolute-energy differences.
Review reproducibility, assets and limitations
Preserve software versions, configuration files, dataset splits, preprocessing choices and checkpoint identifiers with any eventual results. Review code, pretrained weights and training data separately for provenance and applicable permissions; repository metadata alone should not determine asset-use decisions.
Migration and platform constraints also matter. MACE describes a rewrite requiring attention to code and checkpoint migration, while NequIP and TorchANI describe compatibility changes for existing users. TorchANI explicitly leaves AMD GPU and Apple MPS support untested in the linked sources. Record these as source-stated conditions, not compatibility conclusions from this collection.
MACE additionally warns that raw MACE-MP DFT energies are not directly comparable with Materials Project energies carrying compatibility corrections. More broadly, installation success, repository popularity and shared architecture do not establish scientific transferability. Average errors should be accompanied by configuration-level failures and separate inspection of out-of-distribution behaviour.
Short evaluation checklist
- Define chemistry, target outputs and acceptance criteria.
- Select identifiable releases and model assets.
- Mark missing coverage or interface evidence unknown.
- Freeze reference calculations, splits and conversion records.
- Verify units and conventions before scoring.
- Preserve failures, provenance and reproducibility details.
This collection awaits editorial review. It provides no performance ranking, first-hand scientific testing or tested integration claim.