Skip to content

GraphMVP

GraphMVP is a molecular representation-learning research implementation that combines 2D topology and 3D geometry during pre-training, then uses 2D graphs for downstream classification and regression.

Catalog updated ·

Overview

GraphMVP provides the research code accompanying the ICLR 2022 paper “Pre-training Molecular Graph Representation with 3D Geometry.” Its workflow uses molecular topology and geometry together during pre-training, while downstream tasks require only 2D topology. It is therefore relevant to investigating whether geometry-informed pre-training can support molecular property prediction without requiring 3D inputs at the downstream stage.

The repository documents GEOM dataset preparation for separate classification and regression workflows. Its classification atom features include atom types and chirality, whereas regression uses more comprehensive OGB features. Experiment scripts cover pre-training and downstream fine-tuning in both workflows. The README also links pre-trained model weights, training logs, and prediction files, providing artifacts to inspect alongside the implementation rather than presenting the project as a hosted prediction service.

The implementation distinguishes GraphMVP from GraphMVP_hybrid, whose variants add 2D self-supervised pretext tasks. It also includes generative, contrastive, and predictive graph self-supervised baselines for comparative experiments. Older artifact names differ from current script terminology, so matching logs and weights to configurations matters. The documented environment is version-specific, and the source excerpts do not establish current compatibility or quantitative predictive performance. A planned TorchDrug integration is described as future work, not a completed capability.

Key Features

  • Pre-training combines 2D molecular topology with 3D geometry, while downstream tasks use 2D topology only.
  • Separate classification and regression workflows include GEOM preprocessing, pre-training, and fine-tuning scripts.
  • Task-specific atom featurization uses atom types and chirality for classification and more comprehensive OGB features for regression.
  • `GraphMVP_hybrid` variants add 2D self-supervised pretext tasks to the GraphMVP workflow.
  • Implemented comparison baselines include EdgePred, AttrMask, GPT-GNN, InfoGraph, ContextPred, GraphLoG, Grover-Contextual, GraphCL, JOAO, and Grover-Motif.
  • The README links downloadable pre-trained weights, training logs, and prediction files.

Use Cases

  • Suggested evaluation: investigate geometry-informed pre-training for a molecular classification task whose downstream inputs contain only 2D topology.
  • Suggested evaluation: assess the regression workflow with OGB atom features on a representative molecular property dataset.
  • Suggested evaluation: compare `GraphMVP`, `GraphMVP_hybrid`, and supplied graph self-supervised baselines using a consistent downstream evaluation protocol.
  • Inspect linked training logs and prediction files to identify experiment configurations and reconcile older artifact names with current script names.

How to Use

  1. Read the repository README and paper to understand the distinction between geometry-informed pre-training and 2D-only downstream tasks. Select classification or regression before preparing data.
  2. Review the README’s environment instructions, including its pinned Python, PyTorch, and graph-library dependencies. Treat these as the documented setup, not evidence of compatibility with newer environments.
  3. Follow the linked dataset instructions, then inspect the appropriate GEOM preparation workflow. Check data paths and the task-specific atom featurization before processing inputs.
  4. Inspect the relevant pre-training scripts listed in the README. Choose GraphMVP, GraphMVP_hybrid, or a supplied baseline, and verify the configuration rather than assuming the variants are interchangeable.
  5. Consult the linked artifacts and downstream fine-tuning scripts. For an intended evaluation, match weights to their configuration and examine predictions against your chosen task’s labels.

Related resources

ChemBERTa provides BERT-like models for chemical SMILES, with RoBERTa masked-language modelling checkpoints and notebooks for pre-training, fine-tuning and molecular property prediction research.

Open sourcePython

Molecular Property Prediction

ChemML

Open Source

ChemML is a Python suite for chemical and materials data analysis, mining, and modeling, with a modular design and documented work on graph neural networks, AutoML, and explainability.

Open sourcePython

Molecular Property Prediction · Materials Discovery

Chemprop is a PyTorch-based toolkit for training and evaluating message-passing neural networks for molecular property prediction, with CLI workflows, Python modules and task-specific notebooks.

Open sourcePython

Molecular Property Prediction

DeepMol

Open Source

DeepMol is a Python toolkit for molecular machine learning, connecting compound standardization and featurization with model training, evaluation, interpretation and reusable pipelines.

Open sourcePython

Molecular Property Prediction

DGL-LifeSci

Open Source

DGL-LifeSci is a DGL/PyTorch toolkit providing graph learning components and examples for chemistry and biology.

Open sourcePython

Molecular Property Prediction

DimeNet

Model

DimeNet provides reference implementations of DimeNet and DimeNet++ for directional message passing on molecular graphs, with training notebooks, test-set prediction workflows and pretrained models.

Python

Molecular Property Prediction · Quantum Chemistry

Related guides