Overview
GEOM is a molecular conformation dataset intended for property prediction and molecular generation research. The project README describes 37 million conformations covering over 450,000 molecules, with energy and statistical-weight annotations. Its repository provides Jupyter tutorials for accessing and analyzing the data rather than a standalone predictive model. Conformer-based model training is directed to the separate Neural Force Field repository.
The data are distributed through a linked data archive. Four MessagePack archives cover drug-like and QM9 molecules in crude and featurized representations, supporting loading across programming languages. For Python workflows, rdkit_folder.tar.gz provides RDKit mol objects containing coordinates and connectivity. The documented workflow moves from downloaded archives to loaded molecular records, visualizations, derived descriptors, or PDB exports. Lightweight per-molecule JSON summaries allow users to select records before loading individual RDKit pickle files, avoiding the need to load every conformer first.
The README also describes an addition of over 16,000 MoleculeNet species, including CREST ensembles and, for the BACE subset, Hessian data and high-accuracy DFT results. Dedicated tutorials cover loading these records and comparing ensembles from different levels of theory. An important distribution limitation is that the MessagePack files are not being updated as new molecules are added to the Python-specific data. Users should therefore choose a representation with coverage requirements in mind and consult the archive README for checksum verification. The documented dependency environment is a source-reported setup, not evidence of compatibility with other environments.
Key Features
- Energy and statistical-weight annotations for 37 million molecular conformations spanning over 450,000 molecules.
- Four language-agnostic MessagePack archives: `drugs_crude.msgpack.tar.gz`, `drugs_featurized.msgpack.tar.gz`, `qm9_crude.msgpack.tar.gz`, and `qm9_featurized.msgpack.tar.gz`.
- Python-specific RDKit `mol` records containing conformer coordinates and molecular connectivity, with tutorials for visualization, descriptor generation, and PDB export.
- Lightweight `rdkit_folder/{drugs,qm9}_summary.json` files for selecting molecules before loading their individual conformer records.
- MoleculeNet additions with CREST ensembles, plus Hessian data and high-accuracy DFT results for the BACE subset.
- Tutorials for loading MoleculeNet records and comparing ensembles produced at different levels of theory.
Use Cases
- Suggested evaluation: use annotated conformer ensembles as training or evaluation data for molecular generation or conformer-based property prediction, consulting the separately linked Neural Force Field training workflow.
- Suggested analysis: select a property-defined molecular subset through the summary JSON files, then visualize conformers or derive RDKit descriptors without first loading the full dataset.
- Suggested evaluation: compare available MoleculeNet ensembles across levels of theory using the comparison tutorial, including the documented BACE-specific results where applicable.
How to Use
- Read the repository README to identify the dataset representations and documented dependency environment. Choose MessagePack for cross-language access or RDKit records for the Python workflow.
- Visit the data archive, select the relevant files, and consult its README for checksum verification. Account for the stated lack of updates to MessagePack files when assessing molecule coverage.
- Follow the MessagePack loading tutorial for extraction and loading, or the RDKit tutorial for Python-specific records.
- For a targeted analysis, inspect the summary JSON files before loading individual RDKit pickle files. Use the RDKit tutorial to explore conformers, generate descriptors, or export PDB files.
- For the added species, follow the MoleculeNet tutorial and ensemble comparison tutorial. Cite the GEOM paper when using the data.