Skip to content

GEOM

GEOM provides 37 million energy- and statistical-weight-annotated molecular conformations for over 450,000 molecules, with MessagePack data, RDKit objects, and loading and analysis tutorials.

Catalog updated ·

Overview

GEOM is a molecular conformation dataset intended for property prediction and molecular generation research. The project README describes 37 million conformations covering over 450,000 molecules, with energy and statistical-weight annotations. Its repository provides Jupyter tutorials for accessing and analyzing the data rather than a standalone predictive model. Conformer-based model training is directed to the separate Neural Force Field repository.

The data are distributed through a linked data archive. Four MessagePack archives cover drug-like and QM9 molecules in crude and featurized representations, supporting loading across programming languages. For Python workflows, rdkit_folder.tar.gz provides RDKit mol objects containing coordinates and connectivity. The documented workflow moves from downloaded archives to loaded molecular records, visualizations, derived descriptors, or PDB exports. Lightweight per-molecule JSON summaries allow users to select records before loading individual RDKit pickle files, avoiding the need to load every conformer first.

The README also describes an addition of over 16,000 MoleculeNet species, including CREST ensembles and, for the BACE subset, Hessian data and high-accuracy DFT results. Dedicated tutorials cover loading these records and comparing ensembles from different levels of theory. An important distribution limitation is that the MessagePack files are not being updated as new molecules are added to the Python-specific data. Users should therefore choose a representation with coverage requirements in mind and consult the archive README for checksum verification. The documented dependency environment is a source-reported setup, not evidence of compatibility with other environments.

Key Features

  • Energy and statistical-weight annotations for 37 million molecular conformations spanning over 450,000 molecules.
  • Four language-agnostic MessagePack archives: `drugs_crude.msgpack.tar.gz`, `drugs_featurized.msgpack.tar.gz`, `qm9_crude.msgpack.tar.gz`, and `qm9_featurized.msgpack.tar.gz`.
  • Python-specific RDKit `mol` records containing conformer coordinates and molecular connectivity, with tutorials for visualization, descriptor generation, and PDB export.
  • Lightweight `rdkit_folder/{drugs,qm9}_summary.json` files for selecting molecules before loading their individual conformer records.
  • MoleculeNet additions with CREST ensembles, plus Hessian data and high-accuracy DFT results for the BACE subset.
  • Tutorials for loading MoleculeNet records and comparing ensembles produced at different levels of theory.

Use Cases

  • Suggested evaluation: use annotated conformer ensembles as training or evaluation data for molecular generation or conformer-based property prediction, consulting the separately linked Neural Force Field training workflow.
  • Suggested analysis: select a property-defined molecular subset through the summary JSON files, then visualize conformers or derive RDKit descriptors without first loading the full dataset.
  • Suggested evaluation: compare available MoleculeNet ensembles across levels of theory using the comparison tutorial, including the documented BACE-specific results where applicable.

How to Use

  1. Read the repository README to identify the dataset representations and documented dependency environment. Choose MessagePack for cross-language access or RDKit records for the Python workflow.
  2. Visit the data archive, select the relevant files, and consult its README for checksum verification. Account for the stated lack of updates to MessagePack files when assessing molecule coverage.
  3. Follow the MessagePack loading tutorial for extraction and loading, or the RDKit tutorial for Python-specific records.
  4. For a targeted analysis, inspect the summary JSON files before loading individual RDKit pickle files. Use the RDKit tutorial to explore conformers, generate descriptors, or export PDB files.
  5. For the added species, follow the MoleculeNet tutorial and ensemble comparison tutorial. Cite the GEOM paper when using the data.

Related resources

GeoDiff

Model

GeoDiff is a geometric diffusion model for molecular conformation generation, with official code for GEOM-based training, checkpoint sampling, and conformation and property evaluation.

Open sourcePython

Molecular Generation · Computational Chemistry

MolSimplify

Open Source

molSimplify generates inorganic coordination and intermolecular complexes for computational screening, with bundled neural networks for selected properties of octahedral transition metal complexes.

Open sourcePython

Molecular Generation · Computational Chemistry

AIDDISON Explorer is a hosted drug-discovery platform that generates and ranks molecular candidates against target profiles, design constraints, predicted properties and synthetic feasibility.

Molecular Generation · Drug Discovery

Allegro

Model

Allegro implements an E(3)-equivariant interatomic potential as a NequIP extension, with documented GPU acceleration options and a separate plugin for LAMMPS simulations.

Open sourcePython

Computational Chemistry · Materials Discovery

ASE

Open Source

ASE is a Python atomistic simulation library connecting conventional codes and external machine learning potentials through a common calculator interface.

Open sourcePython

Computational Chemistry

An ASE routing Skill in the computational-chemistry-agent-skills collection that separates workflow preparation from calculator configuration and delegates execution elsewhere.

Computational Chemistry · Materials Discovery

Related guides