Overview
Chemprop provides a framework for building molecular property predictors rather than a single pretrained model or hosted prediction service. Its official documentation describes PyTorch-based message-passing neural networks, with command-line tutorials and Python modules supporting training, evaluation and prediction. It fits into research workflows where users develop models for their own chemical property tasks and then apply those models to additional inputs.
The documented data workflow includes datapoints, datasets, dataloaders, splitting and graph featurization for molecules and reactions. Examples cover classification, multicomponent regression, reaction regression and multitask models. Depending on the selected workflow, outputs include property predictions, atom and bond predictions, or learned fingerprint representations. The documentation also lists model saving and loading, ensembling, input/output scaling, extra descriptors and hyperparameter optimization using RayTune or Optuna. Exact input schemas and output formats should be checked in the relevant tutorial; they are not specified in the source excerpts.
Additional notebooks address uncertainty quantification, active learning, transfer learning and interpretation using Myerson values, Monte Carlo tree search and Shapley analysis. These are documented workflow options, not evidence of accuracy on a new dataset. Version selection matters: the README describes a substantial v2 rewrite, changed defaults and discontinued v1 support. It also identifies an edge-update implementation difference from the cited theory papers. The linked sources do not establish predictive performance or suitability for a particular scientific endpoint.
Key Features
- Command-line training and prediction tutorials, alongside Python module tutorials for message-passing neural network models.
- Molecule and reaction graph featurization, with classification, multicomponent regression, reaction regression and multitask examples.
- Atom and bond property prediction, including a constrained prediction notebook.
- Learned fingerprint extraction, model saving and loading, ensembling, and input/output scaling workflows.
- Hyperparameter optimization tutorials and notebooks using RayTune or Optuna.
- Notebook examples for uncertainty quantification, transfer learning and prediction interpretation using Myerson values, Monte Carlo tree search and Shapley analysis.
Use Cases
- Suggested evaluation: train a classification or regression model on a molecular property dataset and assess predictions on a held-out split.
- Suggested evaluation: explore reaction-property or multicomponent regression using the corresponding training and prediction notebooks.
- Suggested evaluation: extract learned molecular fingerprints and assess their usefulness within a separate downstream analysis.
- Suggested evaluation: investigate atom/bond predictions or interpretation methods for a relevant chemical endpoint, checking their usefulness against domain-specific evidence.
How to Use
- Start with the Quickstart and Installation pages. Select the intended release and check its documented environment requirements rather than assuming older setup instructions apply.
- Review the datapoints tutorial and data splitting tutorial. Match your inputs and targets to the chosen task, and plan a separate evaluation split.
- Follow the CLI training tutorial or choose an appropriate notebook example. Begin with a small evaluation run before expanding the workflow.
- Consult saving and loading models and the prediction tutorial to reuse a trained model and inspect its predictions.
- If needed, explore hyperparameter optimization or uncertainty quantification. Record configuration choices and evaluate their effect on your task; these are suggested checks, not reported benchmark results.