Skip to content

SPICE

SPICE provides quantum-mechanical energies, forces and other molecular properties for training machine learning potentials, with an emphasis on drug-like molecules and protein interactions.

Catalog updated ·

Overview

SPICE (Small-Molecule/Protein Interaction Chemical Energies) is a quantum-chemical dataset intended to support training molecular potential functions, particularly for simulations involving drug-like small molecules and proteins. The GitHub repository contains scripts and supporting files used to create the dataset; the dataset itself is distributed separately through Zenodo. It is a training-data resource, not a predictive model or a ready-to-run simulation service.

The project README describes version 2.0 as containing 2,008,628 conformations across 113,999 molecules or clusters, with structures ranging from 2 to 110 atoms. Coverage spans 17 elements, neutral and charged molecules, and low- and high-energy conformations. Subsets include dipeptides, solvated amino acids, PubChem molecules, solvated PubChem molecules, DES370K monomers and dimers, amino acid–ligand pairs, ion pairs and water clusters. These collections target different covalent, non-covalent and solvation interactions rather than a single molecular class.

For molecular configurations, SPICE supplies quantum-mechanical reference results including energies, forces, bond orders, partial charges and atomic multipoles. Calculations use Psi4 at the ωB97M-D3BJ/def2-TZVPPD level of theory. The README provides a sample input for generating additional compatible reference data and stresses that matching the theoretical method alone is insufficient: the program and calculation settings must also match. Versioned releases support reproducible dataset selection, but the source excerpts do not establish trained-model accuracy or generalization. Dataset usage is described as CC0, separately from the repository code’s MIT license.

Key Features

  • Quantum-mechanical reference energies and forces for training molecular potential functions.
  • Additional reference properties including bond orders, partial charges and atomic multipoles.
  • Coverage of 17 elements, charged and neutral molecules, and both low- and high-energy conformations.
  • Distinct subsets for peptide chemistry, drug-like molecules, solvation, protein–ligand contacts, ion interactions and water clusters.
  • Psi4 calculations at the ωB97M-D3BJ/def2-TZVPPD level, with a sample input documenting settings for compatible new calculations.
  • Versioned dataset releases distributed through Zenodo, separate from the repository’s data-generation scripts and supporting files.

Use Cases

  • Intended evaluation: train an energy-and-force potential on selected drug-like molecule subsets and assess it on held-out configurations.
  • Intended evaluation: compare model behavior across peptide, protein–ligand, solvation and ion-pair subsets to examine coverage of different interaction types.
  • Intended evaluation: explore learning partial charges, bond orders or atomic multipoles alongside energies using the documented reference properties.
  • Generate additional reference calculations using the supplied Psi4 settings when extending a SPICE-based training collection.

How to Use

  1. Read the official README to identify the documented subsets, chemical coverage and reference calculation method. Select subsets relevant to your intended molecular system rather than assuming all interactions are equally represented.
  2. Obtain the dataset through the Zenodo DOI. The GitHub repository is not the dataset download; record the specific release used for reproducibility.
  3. Inspect the downloaded release’s files and available properties before designing a training pipeline. For an intended evaluation, define held-out configurations and decide which reference targets your model will learn.
  4. If adding calculations, consult sample.dat in the repository. Match Psi4, ωB97M-D3BJ/def2-TZVPPD and the provided settings; the README cautions against directly combining energies from different programs.
  5. Follow the README’s citation guidance for the dataset papers and version-specific DOI. Distinguish its CC0 data terms from the repository’s MIT code license.

Related resources

QCArchive supports running, storing and sharing quantum chemistry calculations through a database-backed server, a Python client and computation workers maintained in one repository.

Open sourcePython

Quantum Chemistry · Scientific Data

ANI-1

Dataset

ANI-1 provides calculated off-equilibrium molecular conformations, with Python readers for accessing HDF5 files containing coordinates and energies for organic molecules.

Open sourcePython

Quantum Chemistry

Benchling MCP (longevity-genie) is a Python MCP server that connects AI clients to Benchling notebook entries, biological sequences, projects, and entity search using API credentials.

Open sourcePython

Lab Automation · Scientific Data

ChEMBL MCP (cyanheads) connects MCP clients to ChEMBL compound, target, bioactivity and drug records, with structure searches and optional DuckDB analysis of larger activity sets.

Open sourceTypeScript

Drug Discovery · Scientific Data

Official Python client maintained by the ChEMBL group for querying ChEMBL data and cheminformatics services, with QuerySet-style filters, lazy retrieval and local result caching.

Open sourcePython

Drug Discovery · Scientific Data

DeePKS

Model

DeePKS-kit is a Python toolkit for training quantum-chemistry energy functionals, testing post-HF models, and running self-consistent calculations through the DeePHF and DeePKS schemes.

Open sourcePython

Quantum Chemistry

Related guides