Overview
Mordred calculates molecular descriptors for use in cheminformatics workflows. It provides a command-line interface for molecular files and a Python library interface that accepts RDKit molecule objects. Its role is to turn molecular structures into numerical features; the linked sources describe descriptor calculation, not a trained molecular-property prediction model. The README reports 1,613 2D descriptors and 213 3D descriptors, totaling 1,826 in its illustrated configuration.
The command-line workflow accepts SMI, SDF and MOL files, with automatic or explicit input-format selection. Results can be written as CSV-style output to standard output or saved to a file. Users can select descriptor families, supply multiple input files, configure the number of processes, or stream input. In Python, a Calculator can calculate descriptors for an individual RDKit molecule or produce a pandas table from multiple molecules. Documented families include atom and bond counts, ring counts, topological indices, hydrogen-bond descriptors and SLogP.
Configuration matters when preparing comparable feature tables. The library example excludes 3D descriptors with ignore_3D=True, while the command-line 3D option requires SDF or MOL input. Streaming is documented as a low-memory mode without molecule-count information. The README recommends a conda installation route and also describes pip installation after installing RDKit; the supplied setup file declares dependency constraints. These excerpts do not establish compatibility with current environments or predictive accuracy, so installation checks and downstream model evaluation remain separate tasks.
Key Features
- Calculates 2D and 3D molecular descriptors; the README example reports 1,613 2D and 213 3D descriptors.
- Accepts SMI, SDF and MOL files through a command-line interface and writes descriptor results to standard output or a CSV file.
- Supports descriptor-family selection, including ABCIndex and AcidBase, rather than requiring the full descriptor set.
- Processes multiple input files and offers configurable process counts and streaming input.
- Provides a Python Calculator interface for RDKit molecule objects, including single-molecule results and pandas table output.
- Allows 3D descriptors to be excluded in the library interface; command-line 3D calculation requires SDF or MOL input.
Use Cases
- Suggested evaluation: prepare 2D descriptor tables for a QSAR or molecular-property modeling pipeline, then assess model performance separately.
- Suggested evaluation: compare documented structural descriptors across a compound collection using selected families such as RingCount, HydrogenBond or AcidBase.
- Suggested evaluation: assess streaming input for larger molecular-file workflows, checking output completeness and memory use in the intended environment.
- Suggested evaluation: compare 2D-only and 3D-inclusive feature sets using suitable SDF or MOL inputs, without assuming a predictive benefit.
How to Use
- Consult the README for the documented installation routes. It recommends conda; its pip route requires RDKit first. Check the supplied dependency constraints against your intended environment rather than assuming current compatibility.
- Follow the README’s installation-test procedure before processing a collection. Treat this as a local setup check, not validation of scientific suitability.
- Prepare a small representative SMI, SDF or MOL input. Use SDF or MOL when evaluating the command-line 3D option, and decide whether your initial table should contain only 2D descriptors.
- Follow the official examples to choose command-line file output or the RDKit-based Calculator interface. Select descriptor families if the full set is unnecessary; the library also supports pandas output.
- Review the resulting table before downstream use. Consult the master documentation for further guidance, then evaluate descriptor suitability and any predictive model separately on your own data.