Overview
ANI-1 is a molecular reference dataset accompanied by a support repository containing scripts for accessing its records. The README cites a dataset publication describing 20 million calculated off-equilibrium conformations for organic molecules, alongside the original ANI-1 neural network potential paper. This entry concerns the dataset and its extraction utilities, not a runnable potential or a model-training framework.
The downloadable archive is linked through Figshare and expands into an ANI-1_release directory. Its data are distributed across eight HDF5 files named ani_gdb_s0x.h5, where x ranges from 1 to 8 and indicates the number of heavy atoms—C, N and O—in the molecules in that file. The documented quantities include molecular coordinates in Angstroms and energies in Hartrees. The README also supplies self-interaction atomic energy values for H, C, N and O.
The access workflow uses Python, NumPy and h5py. The supplied pyanitools.py includes an anidataloader class for loading and parsing the dataset, while example_data_sampler.py demonstrates sampling through that loader. The archive also includes classes described as supporting loading and storing data in the authors’ in-house format. These utilities provide a starting point for extracting reference data for downstream molecular modelling workflows.
The README specifies Python 3.5 or later and limits its stated reader testing to that version range; it does not establish compatibility with a particular current environment. The source excerpts do not document a complete record schema, training procedure or predictive performance evaluation. The repository’s MIT licence covers its software; dataset reuse terms are not established by these excerpts.
Key Features
- Eight HDF5 data files organize molecules by heavy-atom count, from one to eight C, N or O atoms.
- The `anidataloader` class in `pyanitools.py` loads and parses ANI-1 data.
- `example_data_sampler.py` demonstrates sampling records through the supplied loader.
- The archive includes Python classes for loading and storing data in the authors’ in-house format.
- The README specifies coordinate units in Angstroms, energy units in Hartrees and self-interaction atomic energy values for H, C, N and O.
Use Cases
- Suggested evaluation: extract coordinate–energy samples to assess ANI-1 as reference data for a molecular energy prediction workflow.
- Suggested evaluation: compare data subsets grouped by heavy-atom count when designing training and validation partitions.
- Suggested evaluation: inspect sampled off-equilibrium conformations for suitability in a study of molecular geometry–energy relationships.
How to Use
- Read the support README to understand the archive structure, dependencies and units. Consult the cited dataset paper for scientific context and retain both requested citations when using ANI-1.
- Obtain the dataset archive from the linked Figshare collection. Extract it using the README’s Unix-based archive procedure and locate the resulting
ANI-1_releasedirectory. - Prepare an environment with Python, NumPy and h5py. The README specifies Python 3.5 or later; check suitability for your own environment rather than treating that statement as current compatibility validation.
- Follow the documented setup by adding
ANI-1_release/readers/lib/toPYTHONPATH. Inspectpyanitools.pyand the example reader before runningexample_data_sampler.pyas the documented access check. - Select the relevant
ani_gdb_s0x.h5files by heavy-atom count. Before downstream analysis, verify sampled records, preserve the documented units and check dataset reuse terms separately from the repository’s software licence.