Skip to content

Molfeat

Emmanuel Noutahi

Molfeat is a Python toolkit for small-molecule featurization, combining molecular fingerprints, descriptors and pretrained embeddings with parallel transformation, caching and extensible plugins.

Catalog updated ·

Overview

Molfeat provides a shared interface for turning small molecules into representations for downstream machine-learning workflows. It brings handcrafted featurizers and pretrained molecular embeddings into one package rather than supplying a single property-prediction model. The project README describes a small-molecule-focused 1.x development line, with PyTorch Geometric as its maintained graph backend and integrations including CheMeleon and Mol-JEPA.

The API example starts from SMILES strings sampled from Datamol's FreeSolv dataset. It applies an ECFP calculator to one molecule, then wraps that calculator in MoleculeTransformer for parallel processing of a collection. Outputs are molecular features or embeddings determined by the selected featurizer. Transformer configurations can be saved to and restored from YAML files. ModelStore exposes available models, supports searching by name and provides usage guidance through model cards.

Plugins allow independently developed featurizers to extend the package. Optional dependencies enable functionality such as Transformers integrations, Mordred descriptors, HDF5 and Parquet caching, and S3 or Google Cloud model stores. These dependencies are not all included in the core installation; the README says missing requirements produce an error with installation guidance.

Scope and licensing matter when choosing a workflow. The described 1.x line removes protein featurizers and the legacy DGL, DGLLife and Graphormer adapters. Repository code licensing does not determine the terms for downloaded weights or custom Hub code; Mol-JEPA requires explicit acceptance of a non-commercial licence. The source excerpts do not establish downstream predictive accuracy or comparative runtime results.

Key Features

  • Combines handcrafted molecular featurizers and pretrained embeddings behind a shared package interface.
  • Provides FPCalculator("ecfp") for single-molecule featurization and MoleculeTransformer for parallel processing of molecular collections.
  • Saves and restores MoleculeTransformer configurations through YAML state files.
  • Supports listing and searching models with ModelStore, including model-card usage guidance.
  • Extends featurization through independently developed plugins and provides optional HDF5 and Parquet cache support.
  • Describes maintained CheMeleon and Mol-JEPA integrations with lazy model loading and explicit external checkpoint licensing.

Use Cases

  • Suggested evaluation: compare ECFP features and pretrained embeddings as inputs to a small-molecule property-prediction pipeline using a fixed dataset and evaluation protocol.
  • Suggested evaluation: assess whether saved transformer configurations and optional caching support repeatable featurization across repeated dataset-processing runs.
  • Suggested evaluation: integrate a custom molecular featurizer through the plugin system and check its outputs against the project's intended representation requirements.

How to Use

  1. Read the official documentation and repository README to choose a featurizer. Distinguish the described 1.x development requirements from the release you intend to install.
  2. Follow the README's installation instructions for uv, pip or conda-forge. Select only the documented optional dependencies needed for your chosen representation, cache format or model store.
  3. Use the README's API tour as a starting point: prepare a small collection of SMILES strings, apply FPCalculator("ecfp") to one molecule, then use MoleculeTransformer for the collection.
  4. For pretrained representations, inspect ModelStore listings and the selected model card's usage guidance. Check checkpoint and custom Hub-code terms separately from the repository code licence before downloading artifacts.
  5. Save the transformer configuration to YAML and restore it as shown in the API tour. As an intended evaluation, inspect output dimensions, failed inputs and repeatability on representative molecules before incorporating features into a downstream model.

Related resources

Mordred

Open Source

Mordred is a Python molecular descriptor calculator with command-line and RDKit-based library workflows, producing descriptor tables for cheminformatics analysis and downstream model evaluation.

Open sourcePython

Molecular Property Prediction · Cheminformatics

ChemBERTa provides BERT-like models for chemical SMILES, with RoBERTa masked-language modelling checkpoints and notebooks for pre-training, fine-tuning and molecular property prediction research.

Open sourcePython

Molecular Property Prediction

ChemCP is an MCP App that turns SMILES strings into interactive 2D molecular diagrams and computed descriptors using RDKit.js, with an in-chat viewer for hosts that support MCP Apps.

TypeScriptJavaScript

Cheminformatics

ChemCrow is a Python chemistry agent package built with Langchain that connects language-model reasoning to RDKit, paper-qa, chemical databases, and reaction-planning tools.

Open sourcePython

Cheminformatics · Literature & Research

Chemistry Development Kit (CDK) is a Java library for chemical representations, file conversion, structure searching, fingerprints, QSAR descriptors and molecular rendering within other programs.

Open source

Cheminformatics

ChemML

Open Source

ChemML is a Python suite for chemical and materials data analysis, mining, and modeling, with a modular design and documented work on graph neural networks, AutoML, and explainability.

Open sourcePython

Molecular Property Prediction · Materials Discovery

Related guides

Models

Choosing between Chemprop and DeepChem

Compare Chemprop’s molecular property prediction focus with DeepChem’s broader scientific machine-learning scope, then plan a fair evaluation using shared data, explicit inputs and outputs, and reproducible decision criteria.

Deployment

Local Deployment Checklist for Chemistry AI

Plan an isolated, reproducible local environment for RDKit, Chemprop or DeepChem. Check documented dependencies, data movement, example outputs and recovery procedures before using sensitive chemistry data.