Overview
DeepMol is a Python machine and deep learning framework for drug discovery and computational chemistry. It connects molecular data preparation with predictive modelling, using RDKit for molecular operations and wrappers for Scikit-Learn, Keras and DeepChem models. It is a workflow toolkit rather than a single predictive model: users can assemble preprocessing and modelling components or use the documented pipeline interface to combine them.
Inputs include CSV datasets containing SMILES and optional identifiers, labels or existing features, as well as SDF files containing molecular structures. Preparation options include basic sanitization, configurable standardization and ChEMBL standardization. Molecular representations include several fingerprint families, DeepChem featurizers and Mol2Vec embeddings. The workflow can produce feature matrices, transformed datasets, trained models, predictions and evaluation metrics; models and pipelines can also be saved and loaded.
For model development, the README describes random and stratified splitting, cross-validation, feature selection, imbalance handling, hyperparameter search and pipeline optimization. Exploration methods include PCA, tSNE, KMeans and UMAP. SHAP-based explanations and fingerprint-bit drawings support inspection of feature contributions. Separately linked case-study repositories provide deployed models, including ADMET examples; these are distinct from the core toolkit.
The README flags dependency constraints: GPU use requires TensorFlow and DGL versions matched to the hardware’s CUDA drivers, JAX can cause dependency conflicts, and loading TensorFlow models on macOS has a documented issue. Seq2Seq and transformer-based molecular embeddings are described as under development, not available capabilities. The supplied evidence does not establish predictive performance on a user’s dataset.
Key Features
- Loads molecular datasets from CSV with SMILES and optional labels, identifiers or features, and from SDF files with molecular structures.
- Provides basic, configurable and ChEMBL standardization, including options for charge neutralization, isotope removal, stereochemistry removal and fragment selection.
- Computes Morgan, MACCS, Layered, RDK and AtomPair fingerprints, alongside DeepChem featurizers and Mol2Vec embeddings.
- Wraps Scikit-Learn, Keras and DeepChem models for training, prediction, metric-based evaluation, cross-validation and model persistence.
- Supports feature selection, random and stratified data splits, imbalance sampling, hyperparameter searches and optimization of multi-step pipelines.
- Provides SHAP explanations, fingerprint-bit visualization and unsupervised exploration through PCA, tSNE, KMeans and UMAP.
Use Cases
- Suggested evaluation: build a molecular-property classification or regression workflow from labelled SMILES, comparing representations and models on held-out data.
- Suggested evaluation: assess how standardization choices and imbalance sampling affect a molecular activity-classification workflow.
- Suggested evaluation: inspect SHAP contributions and corresponding fingerprint bits to investigate which structural features influence predictions.
- Suggested evaluation: explore the separately linked ADMET case-study models for predictions on new compounds, checking each model’s requirements and applicability before use.
How to Use
- Start with the DeepMol documentation and the repository README. Select the documented installation route and dependencies needed for preprocessing, machine learning or deep learning; review the CUDA, JAX and macOS cautions before configuring the environment.
- Prepare a small representative dataset. Follow the README’s CSVLoader example for SMILES with optional labels and identifiers, or its SDFLoader example for structure files. Check the resulting dataset dimensions and label fields.
- Choose a standardization policy and molecular representation from the documented options. Inspect the transformed molecules and feature dimensions before extending the workflow to the full dataset.
- Define training, validation and test partitions, then select a wrapped model or assemble a pipeline. For guided task examples, consult the README’s linked classification, regression or multi-task AutoML notebooks.
- Evaluate predictions using task-appropriate metrics, inspect SHAP or fingerprint explanations where relevant, and save the model or pipeline. If using deployed models instead, consult the separate case-study repository rather than treating the toolkit itself as a pretrained predictor.