Skip to content

Chemprop

The Chemprop Development Team

Chemprop is a PyTorch-based toolkit for training and evaluating message-passing neural networks for molecular property prediction, with CLI workflows, Python modules and task-specific notebooks.

Catalog updated ·

Overview

Chemprop provides a framework for building molecular property predictors rather than a single pretrained model or hosted prediction service. Its official documentation describes PyTorch-based message-passing neural networks, with command-line tutorials and Python modules supporting training, evaluation and prediction. It fits into research workflows where users develop models for their own chemical property tasks and then apply those models to additional inputs.

The documented data workflow includes datapoints, datasets, dataloaders, splitting and graph featurization for molecules and reactions. Examples cover classification, multicomponent regression, reaction regression and multitask models. Depending on the selected workflow, outputs include property predictions, atom and bond predictions, or learned fingerprint representations. The documentation also lists model saving and loading, ensembling, input/output scaling, extra descriptors and hyperparameter optimization using RayTune or Optuna. Exact input schemas and output formats should be checked in the relevant tutorial; they are not specified in the source excerpts.

Additional notebooks address uncertainty quantification, active learning, transfer learning and interpretation using Myerson values, Monte Carlo tree search and Shapley analysis. These are documented workflow options, not evidence of accuracy on a new dataset. Version selection matters: the README describes a substantial v2 rewrite, changed defaults and discontinued v1 support. It also identifies an edge-update implementation difference from the cited theory papers. The linked sources do not establish predictive performance or suitability for a particular scientific endpoint.

Key Features

  • Command-line training and prediction tutorials, alongside Python module tutorials for message-passing neural network models.
  • Molecule and reaction graph featurization, with classification, multicomponent regression, reaction regression and multitask examples.
  • Atom and bond property prediction, including a constrained prediction notebook.
  • Learned fingerprint extraction, model saving and loading, ensembling, and input/output scaling workflows.
  • Hyperparameter optimization tutorials and notebooks using RayTune or Optuna.
  • Notebook examples for uncertainty quantification, transfer learning and prediction interpretation using Myerson values, Monte Carlo tree search and Shapley analysis.

Use Cases

  • Suggested evaluation: train a classification or regression model on a molecular property dataset and assess predictions on a held-out split.
  • Suggested evaluation: explore reaction-property or multicomponent regression using the corresponding training and prediction notebooks.
  • Suggested evaluation: extract learned molecular fingerprints and assess their usefulness within a separate downstream analysis.
  • Suggested evaluation: investigate atom/bond predictions or interpretation methods for a relevant chemical endpoint, checking their usefulness against domain-specific evidence.

How to Use

  1. Start with the Quickstart and Installation pages. Select the intended release and check its documented environment requirements rather than assuming older setup instructions apply.
  2. Review the datapoints tutorial and data splitting tutorial. Match your inputs and targets to the chosen task, and plan a separate evaluation split.
  3. Follow the CLI training tutorial or choose an appropriate notebook example. Begin with a small evaluation run before expanding the workflow.
  4. Consult saving and loading models and the prediction tutorial to reuse a trained model and inspect its predictions.
  5. If needed, explore hyperparameter optimization or uncertainty quantification. Record configuration choices and evaluate their effect on your task; these are suggested checks, not reported benchmark results.

Related resources

ChemBERTa provides BERT-like models for chemical SMILES, with RoBERTa masked-language modelling checkpoints and notebooks for pre-training, fine-tuning and molecular property prediction research.

Open sourcePython

Molecular Property Prediction

ChemML

Open Source

ChemML is a Python suite for chemical and materials data analysis, mining, and modeling, with a modular design and documented work on graph neural networks, AutoML, and explainability.

Open sourcePython

Molecular Property Prediction · Materials Discovery

DeepMol

Open Source

DeepMol is a Python toolkit for molecular machine learning, connecting compound standardization and featurization with model training, evaluation, interpretation and reusable pipelines.

Open sourcePython

Molecular Property Prediction

DGL-LifeSci

Open Source

DGL-LifeSci is a DGL/PyTorch toolkit providing graph learning components and examples for chemistry and biology.

Open sourcePython

Molecular Property Prediction

DimeNet

Model

DimeNet provides reference implementations of DimeNet and DimeNet++ for directional message passing on molecular graphs, with training notebooks, test-set prediction workflows and pretrained models.

Python

Molecular Property Prediction · Quantum Chemistry

DrugAgent is a research multi-agent framework that combines LLM planning with domain-guided code generation to build and evaluate machine-learning workflows for drug discovery.

Molecular Property Prediction · Drug Discovery

Related guides

Overview

How to Build a Chemistry AI Agent Stack

Design a chemistry AI agent stack by separating language models, orchestration, Skills, MCP interfaces, APIs, libraries and data. Follow proposed molecular and materials workflows, define input/output contracts, and plan permissions, evaluation and reproducible deployment.

Models

Choosing between Chemprop and DeepChem

Compare Chemprop’s molecular property prediction focus with DeepChem’s broader scientific machine-learning scope, then plan a fair evaluation using shared data, explicit inputs and outputs, and reproducible decision criteria.

Deployment

Local Deployment Checklist for Chemistry AI

Plan an isolated, reproducible local environment for RDKit, Chemprop or DeepChem. Check documented dependencies, data movement, example outputs and recovery procedures before using sensitive chemistry data.