Skip to content

MolCLR

Yuyang Wang / Carnegie Mellon University

MolCLR is a molecular contrastive-learning framework that pre-trains graph neural networks on unlabelled molecules and supports fine-tuning for downstream molecular property prediction.

Catalog updated ·

Overview

MolCLR is the official implementation accompanying “Molecular Contrastive Learning of Representations via Graph Neural Networks.” It provides a research workflow for learning molecular representations from unlabelled data before adapting graph neural networks to molecular property prediction tasks. The README describes pre-training on approximately 10 million unique molecules, placing the resource between molecular dataset preparation and supervised downstream modelling rather than presenting it as a ready-to-use prediction service.

The documented workflow has separate pre-training and fine-tuning stages. Pre-training uses the supplied molecular dataset, identified as pubchem-10m-clean.txt, with settings described in config.yaml. Fine-tuning uses benchmark datasets stored in folders named after each benchmark, with its own configuration in config_finetune.yaml. The repository also supplies pre-trained GCN and GIN models, allowing users to explore downstream adaptation without first repeating the full pre-training stage. Its documented outputs include trained model checkpoints and training information monitored through TensorBoard.

The README links the datasets used in the paper and points to MoleculeNet as another source of benchmarks. It also gives a version-pinned Python environment recipe using PyTorch, PyTorch Geometric and RDKit. These instructions document the original setup, not verified compatibility with current systems. The source excerpts do not specify prediction output formats, deployment interfaces or quantitative benchmark results; suitability for a new property dataset therefore remains an evaluation question.

Key Features

  • Contrastive pre-training of graph neural networks on a large unlabelled molecular dataset described as approximately 10 million unique molecules.
  • Separate configurable pre-training and downstream fine-tuning workflows through `config.yaml` and `config_finetune.yaml`.
  • Pre-trained GCN and GIN models supplied in the repository for downstream adaptation.
  • Downloadable pre-training data and paper benchmark datasets, with a documented `./data` directory layout.
  • TensorBoard monitoring of training logs under checkpoint paths.

Use Cases

  • Intended evaluation: fine-tune a supplied pre-trained model on a molecular property benchmark and assess its suitability for the selected prediction task.
  • Intended evaluation: compare models adapted from MolCLR checkpoints with an independently defined baseline using the same dataset and evaluation protocol.
  • Intended evaluation: investigate molecular representation learning by changing documented pre-training configuration settings and examining subsequent fine-tuning outcomes.

How to Use

  1. Read the official README and consult the paper to distinguish the representation-learning stage from downstream property modelling.
  2. Obtain the repository and follow its documented environment recipe. The README specifies Python 3.7 and pinned dependencies; treat these as source instructions rather than evidence of current compatibility.
  3. Download the supplied datasets and extract them under ./data. Identify pubchem-10m-clean.txt for pre-training and the appropriate benchmark folder for fine-tuning.
  4. Inspect config.yaml before using the documented molclr.py entry point. If repeating pre-training, follow the README’s TensorBoard guidance to inspect training logs.
  5. For downstream evaluation, inspect config_finetune.yaml, choose a supplied pre-trained model and use the documented finetune.py workflow. Define your evaluation protocol separately; the excerpt does not establish performance on your dataset.

Related resources

ChemBERTa provides BERT-like models for chemical SMILES, with RoBERTa masked-language modelling checkpoints and notebooks for pre-training, fine-tuning and molecular property prediction research.

Open sourcePython

Molecular Property Prediction

ChemML

Open Source

ChemML is a Python suite for chemical and materials data analysis, mining, and modeling, with a modular design and documented work on graph neural networks, AutoML, and explainability.

Open sourcePython

Molecular Property Prediction · Materials Discovery

Chemprop is a PyTorch-based toolkit for training and evaluating message-passing neural networks for molecular property prediction, with CLI workflows, Python modules and task-specific notebooks.

Open sourcePython

Molecular Property Prediction

DeepMol

Open Source

DeepMol is a Python toolkit for molecular machine learning, connecting compound standardization and featurization with model training, evaluation, interpretation and reusable pipelines.

Open sourcePython

Molecular Property Prediction

DGL-LifeSci

Open Source

DGL-LifeSci is a DGL/PyTorch toolkit providing graph learning components and examples for chemistry and biology.

Open sourcePython

Molecular Property Prediction

DimeNet

Model

DimeNet provides reference implementations of DimeNet and DimeNet++ for directional message passing on molecular graphs, with training notebooks, test-set prediction workflows and pretrained models.

Python

Molecular Property Prediction · Quantum Chemistry

Related guides