Overview
Molfeat provides a shared interface for turning small molecules into representations for downstream machine-learning workflows. It brings handcrafted featurizers and pretrained molecular embeddings into one package rather than supplying a single property-prediction model. The project README describes a small-molecule-focused 1.x development line, with PyTorch Geometric as its maintained graph backend and integrations including CheMeleon and Mol-JEPA.
The API example starts from SMILES strings sampled from Datamol's FreeSolv dataset. It applies an ECFP calculator to one molecule, then wraps that calculator in MoleculeTransformer for parallel processing of a collection. Outputs are molecular features or embeddings determined by the selected featurizer. Transformer configurations can be saved to and restored from YAML files. ModelStore exposes available models, supports searching by name and provides usage guidance through model cards.
Plugins allow independently developed featurizers to extend the package. Optional dependencies enable functionality such as Transformers integrations, Mordred descriptors, HDF5 and Parquet caching, and S3 or Google Cloud model stores. These dependencies are not all included in the core installation; the README says missing requirements produce an error with installation guidance.
Scope and licensing matter when choosing a workflow. The described 1.x line removes protein featurizers and the legacy DGL, DGLLife and Graphormer adapters. Repository code licensing does not determine the terms for downloaded weights or custom Hub code; Mol-JEPA requires explicit acceptance of a non-commercial licence. The source excerpts do not establish downstream predictive accuracy or comparative runtime results.
Key Features
- Combines handcrafted molecular featurizers and pretrained embeddings behind a shared package interface.
- Provides FPCalculator("ecfp") for single-molecule featurization and MoleculeTransformer for parallel processing of molecular collections.
- Saves and restores MoleculeTransformer configurations through YAML state files.
- Supports listing and searching models with ModelStore, including model-card usage guidance.
- Extends featurization through independently developed plugins and provides optional HDF5 and Parquet cache support.
- Describes maintained CheMeleon and Mol-JEPA integrations with lazy model loading and explicit external checkpoint licensing.
Use Cases
- Suggested evaluation: compare ECFP features and pretrained embeddings as inputs to a small-molecule property-prediction pipeline using a fixed dataset and evaluation protocol.
- Suggested evaluation: assess whether saved transformer configurations and optional caching support repeatable featurization across repeated dataset-processing runs.
- Suggested evaluation: integrate a custom molecular featurizer through the plugin system and check its outputs against the project's intended representation requirements.
How to Use
- Read the official documentation and repository README to choose a featurizer. Distinguish the described 1.x development requirements from the release you intend to install.
- Follow the README's installation instructions for uv, pip or conda-forge. Select only the documented optional dependencies needed for your chosen representation, cache format or model store.
- Use the README's API tour as a starting point: prepare a small collection of SMILES strings, apply FPCalculator("ecfp") to one molecule, then use MoleculeTransformer for the collection.
- For pretrained representations, inspect ModelStore listings and the selected model card's usage guidance. Check checkpoint and custom Hub-code terms separately from the repository code licence before downloading artifacts.
- Save the transformer configuration to YAML and restore it as shown in the API tour. As an intended evaluation, inspect output dimensions, failed inputs and repeatability on representative molecules before incorporating features into a downstream model.