Overview
GraphMVP provides the research code accompanying the ICLR 2022 paper “Pre-training Molecular Graph Representation with 3D Geometry.” Its workflow uses molecular topology and geometry together during pre-training, while downstream tasks require only 2D topology. It is therefore relevant to investigating whether geometry-informed pre-training can support molecular property prediction without requiring 3D inputs at the downstream stage.
The repository documents GEOM dataset preparation for separate classification and regression workflows. Its classification atom features include atom types and chirality, whereas regression uses more comprehensive OGB features. Experiment scripts cover pre-training and downstream fine-tuning in both workflows. The README also links pre-trained model weights, training logs, and prediction files, providing artifacts to inspect alongside the implementation rather than presenting the project as a hosted prediction service.
The implementation distinguishes GraphMVP from GraphMVP_hybrid, whose variants add 2D self-supervised pretext tasks. It also includes generative, contrastive, and predictive graph self-supervised baselines for comparative experiments. Older artifact names differ from current script terminology, so matching logs and weights to configurations matters. The documented environment is version-specific, and the source excerpts do not establish current compatibility or quantitative predictive performance. A planned TorchDrug integration is described as future work, not a completed capability.
Key Features
- Pre-training combines 2D molecular topology with 3D geometry, while downstream tasks use 2D topology only.
- Separate classification and regression workflows include GEOM preprocessing, pre-training, and fine-tuning scripts.
- Task-specific atom featurization uses atom types and chirality for classification and more comprehensive OGB features for regression.
- `GraphMVP_hybrid` variants add 2D self-supervised pretext tasks to the GraphMVP workflow.
- Implemented comparison baselines include EdgePred, AttrMask, GPT-GNN, InfoGraph, ContextPred, GraphLoG, Grover-Contextual, GraphCL, JOAO, and Grover-Motif.
- The README links downloadable pre-trained weights, training logs, and prediction files.
Use Cases
- Suggested evaluation: investigate geometry-informed pre-training for a molecular classification task whose downstream inputs contain only 2D topology.
- Suggested evaluation: assess the regression workflow with OGB atom features on a representative molecular property dataset.
- Suggested evaluation: compare `GraphMVP`, `GraphMVP_hybrid`, and supplied graph self-supervised baselines using a consistent downstream evaluation protocol.
- Inspect linked training logs and prediction files to identify experiment configurations and reconcile older artifact names with current script names.
How to Use
- Read the repository README and paper to understand the distinction between geometry-informed pre-training and 2D-only downstream tasks. Select classification or regression before preparing data.
- Review the README’s environment instructions, including its pinned Python, PyTorch, and graph-library dependencies. Treat these as the documented setup, not evidence of compatibility with newer environments.
- Follow the linked dataset instructions, then inspect the appropriate GEOM preparation workflow. Check data paths and the task-specific atom featurization before processing inputs.
- Inspect the relevant pre-training scripts listed in the README. Choose
GraphMVP,GraphMVP_hybrid, or a supplied baseline, and verify the configuration rather than assuming the variants are interchangeable. - Consult the linked artifacts and downstream fine-tuning scripts. For an intended evaluation, match weights to their configuration and examine predictions against your chosen task’s labels.