Overview
SPICE (Small-Molecule/Protein Interaction Chemical Energies) is a quantum-chemical dataset intended to support training molecular potential functions, particularly for simulations involving drug-like small molecules and proteins. The GitHub repository contains scripts and supporting files used to create the dataset; the dataset itself is distributed separately through Zenodo. It is a training-data resource, not a predictive model or a ready-to-run simulation service.
The project README describes version 2.0 as containing 2,008,628 conformations across 113,999 molecules or clusters, with structures ranging from 2 to 110 atoms. Coverage spans 17 elements, neutral and charged molecules, and low- and high-energy conformations. Subsets include dipeptides, solvated amino acids, PubChem molecules, solvated PubChem molecules, DES370K monomers and dimers, amino acid–ligand pairs, ion pairs and water clusters. These collections target different covalent, non-covalent and solvation interactions rather than a single molecular class.
For molecular configurations, SPICE supplies quantum-mechanical reference results including energies, forces, bond orders, partial charges and atomic multipoles. Calculations use Psi4 at the ωB97M-D3BJ/def2-TZVPPD level of theory. The README provides a sample input for generating additional compatible reference data and stresses that matching the theoretical method alone is insufficient: the program and calculation settings must also match. Versioned releases support reproducible dataset selection, but the source excerpts do not establish trained-model accuracy or generalization. Dataset usage is described as CC0, separately from the repository code’s MIT license.
Key Features
- Quantum-mechanical reference energies and forces for training molecular potential functions.
- Additional reference properties including bond orders, partial charges and atomic multipoles.
- Coverage of 17 elements, charged and neutral molecules, and both low- and high-energy conformations.
- Distinct subsets for peptide chemistry, drug-like molecules, solvation, protein–ligand contacts, ion interactions and water clusters.
- Psi4 calculations at the ωB97M-D3BJ/def2-TZVPPD level, with a sample input documenting settings for compatible new calculations.
- Versioned dataset releases distributed through Zenodo, separate from the repository’s data-generation scripts and supporting files.
Use Cases
- Intended evaluation: train an energy-and-force potential on selected drug-like molecule subsets and assess it on held-out configurations.
- Intended evaluation: compare model behavior across peptide, protein–ligand, solvation and ion-pair subsets to examine coverage of different interaction types.
- Intended evaluation: explore learning partial charges, bond orders or atomic multipoles alongside energies using the documented reference properties.
- Generate additional reference calculations using the supplied Psi4 settings when extending a SPICE-based training collection.
How to Use
- Read the official README to identify the documented subsets, chemical coverage and reference calculation method. Select subsets relevant to your intended molecular system rather than assuming all interactions are equally represented.
- Obtain the dataset through the Zenodo DOI. The GitHub repository is not the dataset download; record the specific release used for reproducibility.
- Inspect the downloaded release’s files and available properties before designing a training pipeline. For an intended evaluation, define held-out configurations and decide which reference targets your model will learn.
- If adding calculations, consult
sample.datin the repository. Match Psi4, ωB97M-D3BJ/def2-TZVPPD and the provided settings; the README cautions against directly combining energies from different programs. - Follow the README’s citation guidance for the dataset papers and version-specific DOI. Distinguish its CC0 data terms from the repository’s MIT code license.