Skip to content

MoFlow

MoFlow is an invertible molecular graph generation model with documented workflows for QM9 and zinc250k training, reconstruction, latent-space sampling, interpolation and property optimization.

Catalog updated ·

Overview

MoFlow is a research model for generating molecular graphs through an invertible flow formulation. The repository accompanies the paper by Chengxi Zang and Fei Wang and provides workflows for QM9 and zinc250k. Its role is molecular generation and latent-space exploration, with additional property-regression models used for optimization. The README covers preprocessing, training and experiments rather than a hosted service or general-purpose chemistry interface.

The documented input workflow starts with SMILES strings, which the preprocessing scripts convert into molecular graphs. Users can train a dataset-specific model or obtain the linked trained models. Generation and reconstruction use a model snapshot together with its hyperparameter configuration. Documented outputs include reconstructed graphs, sampled molecules, molecular figures and experiment logs. Interpolation workflows explore paths between two molecules or a grid in latent space, while random generation exposes sampling temperature, sample count and optional validity correction.

For property optimization, the repository describes training an additional MLP that maps latent representations to QED or plogp. This regression model is then used to generate ranked molecular candidates or optimize existing molecules under a similarity constraint. These are computational property-optimization workflows; the supplied evidence does not establish experimental activity or suitability for a particular discovery program.

The examples are tied to the supplied datasets, model configurations and recorded dependency environment. The README includes author-reported experimental results, but these should not be treated as guaranteed outcomes for other settings. Validity correction is an explicit generation option, so evaluations should distinguish corrected outputs from uncorrected sampling. Compatibility with current software environments and behavior on other datasets remain unestablished by the supplied excerpt.

Key Features

  • Preprocesses SMILES strings into molecular graphs for the documented QM9 and zinc250k workflows.
  • Provides dataset-specific model-training examples and a link to trained model files.
  • Supports molecular reconstruction and random latent-space generation with configurable temperature, sample count and validity correction.
  • Includes two-molecule and grid-based latent-space interpolation workflows with molecular visualization.
  • Trains latent-space MLP regressors for QED or plogp and uses them to produce ranked optimized molecular candidates.
  • Supports constrained property optimization with a configurable molecular similarity cutoff.

Use Cases

  • Intended evaluation: assess reconstruction and generation on QM9 or zinc250k, recording validity, novelty and uniqueness separately for corrected and uncorrected outputs.
  • Intended evaluation: explore structural changes along two-point or grid-based latent-space interpolations using the documented molecular visualizations.
  • Intended evaluation: generate QED-ranked candidates or examine plogp optimization under different similarity constraints, without treating computational scores as experimental validation.

How to Use

  1. Read the repository README and the MoFlow paper to select a documented dataset and experiment: reconstruction, random generation, interpolation or optimization.
  2. Follow the README’s environment and dependency instructions. Treat its recorded package versions as historical setup information, not proof of compatibility with a current environment.
  3. Prepare QM9 or zinc250k through the documented data_preprocess.py workflow, which converts SMILES strings to molecular graphs. Keep the dataset choice consistent across preprocessing, training and generation.
  4. Train with the corresponding train_model.py example or obtain files from the linked trained-model folder. Match the snapshot and moflow-params.json configuration to the selected model.
  5. Use the relevant generate.py workflow. For an intended evaluation, record sampling settings, whether validity correction is enabled, and the resulting logs or molecular figures.
  6. For optimization, follow optimize_property.py to train or select the property regressor. Check the property-model path and similarity settings before interpreting ranked candidates or constrained-optimization outputs.

Related resources

AIDDISON Explorer is a hosted drug-discovery platform that generates and ranks molecular candidates against target profiles, design constraints, predicted properties and synthetic feasibility.

Molecular Generation · Drug Discovery

Chemistry42

Platforms

Chemistry42 is a small-molecule discovery platform combining generative design, retrosynthesis, ADMET and selectivity prediction, and physics-based prioritization for hit identification and lead optimization.

Molecular Generation · Drug Discovery

Official research code for E(3)-equivariant diffusion models that generate 3D molecules, with QM9 and GEOM-Drugs training workflows, sample analysis, and property-conditioned generation.

Open sourcePython

Molecular Generation

FEgrow

Open Source

FEgrow supports interactive ligand-series construction for free-energy preparation, with documented active-learning examples and Dask-based acceleration for molecular design workflows.

Open source

Molecular Generation · Drug Discovery

GeoDiff

Model

GeoDiff is a geometric diffusion model for molecular conformation generation, with official code for GEOM-based training, checkpoint sampling, and conformation and property evaluation.

Open sourcePython

Molecular Generation · Computational Chemistry

GEOM

Dataset

GEOM provides 37 million energy- and statistical-weight-annotated molecular conformations for over 450,000 molecules, with MessagePack data, RDKit objects, and loading and analysis tutorials.

Python

Molecular Generation · Computational Chemistry

Related guides

Workflows

AI for Drug Discovery: Agents, MCP Servers and Platforms

Choose tools by workflow stage: target evidence, protein structures, known binding sites, molecular generation, property analysis and screening. Compare research agents, MCP integrations, hosted APIs and platforms, with a proposed end-to-end example and practical evaluation gates.

Workflows

Reproducibility for Chemistry AI Workflows

Build a compact evaluation record that connects scientific questions to data provenance, transformations, model configurations, outputs and failures, with practical planning examples for chemistry AI.