Overview
ChemML is a machine learning and informatics toolkit for chemical and materials research. It supports analysis, data mining, and modeling rather than providing a single pretrained predictor. The README describes a modular, object-oriented Python design with a library format similar to Scikit-learn, positioning it as a component of researcher-built data workflows.
Its stated inputs are chemical and materials data. The source excerpts do not specify accepted molecular representations, file formats, target-property schemas, or exact output structures. They identify work on graph convolution neural networks, explainable AI, automated model optimization, AutoML, a Jupyter graphical interface, and a port of the Magpie descriptor library. These establish areas of functionality, but not detailed API behavior or demonstrated predictive accuracy.
ChemML draws on scientific Python and domain-specific libraries. Package metadata declares dependencies including RDKit, mordred, TensorFlow, scikit-learn, SHAP, and LIME. The README recommends an Anaconda environment and separately directs users to install PyTorch according to their operating system and GPU configuration. It also lists optional dependencies for the Jupyter wrapper and AutoML screening, so setup requirements can vary with the intended workflow.
The repository code carries a BSD 3-Clause license. The excerpts do not establish separate terms for bundled datasets or model weights, nor do they provide benchmark results or verified platform compatibility. Researchers evaluating ChemML should consult its linked documentation for task-specific inputs and outputs and assess the chosen workflow on representative data.
Key Features
- Python program suite for analysis, mining, and modeling of chemical and materials data.
- Modular, object-oriented library design with an interface style described as similar to Scikit-learn.
- Graph convolution neural-network and explainable-AI functionality identified in the README's contributor descriptions.
- AutoML screening and automated model optimization, with xgboost and lightgbm listed as optional user-installed dependencies.
- Jupyter GUI and wrapper support, with python-graphviz listed as an optional dependency for the wrapper.
- A port of the Magpie descriptor library, with associated lookup data included in the package-data configuration.
Use Cases
- Intended evaluation: assess ChemML as a toolkit for molecular-property modeling on a representative chemical dataset, after checking supported representations and target handling.
- Intended evaluation: investigate Magpie-based descriptors for a materials-data modeling workflow, confirming the documented descriptor inputs before preparing data.
- Intended evaluation: explore AutoML screening or automated model optimization for a chemistry modeling task and compare candidate models using a predefined validation protocol.
- Intended evaluation: examine graph convolution neural-network and explainability functionality for a research workflow where model interpretation is important.
How to Use
- Start with the ChemML documentation to identify the relevant workflow and confirm its expected inputs, outputs, and API. The source excerpts do not define those details.
- Check the release history and changelog before selecting a revision; no particular release is established here.
- Follow the environment and installation guidance in the repository README. It recommends Anaconda and includes Open Babel setup. Use the linked PyTorch installation guide for operating-system and GPU-specific choices.
- Identify optional dependencies for the chosen workflow. The README lists python-graphviz for the Jupyter wrapper and xgboost or lightgbm for AutoML screening; these are separate from the main package dependencies.
- As an intended evaluation, run a documented workflow on a small representative dataset, inspect its outputs, and record the environment and validation approach. Consult the repository code license separately from any dataset or model terms.