Overview
MoFlow is a research model for generating molecular graphs through an invertible flow formulation. The repository accompanies the paper by Chengxi Zang and Fei Wang and provides workflows for QM9 and zinc250k. Its role is molecular generation and latent-space exploration, with additional property-regression models used for optimization. The README covers preprocessing, training and experiments rather than a hosted service or general-purpose chemistry interface.
The documented input workflow starts with SMILES strings, which the preprocessing scripts convert into molecular graphs. Users can train a dataset-specific model or obtain the linked trained models. Generation and reconstruction use a model snapshot together with its hyperparameter configuration. Documented outputs include reconstructed graphs, sampled molecules, molecular figures and experiment logs. Interpolation workflows explore paths between two molecules or a grid in latent space, while random generation exposes sampling temperature, sample count and optional validity correction.
For property optimization, the repository describes training an additional MLP that maps latent representations to QED or plogp. This regression model is then used to generate ranked molecular candidates or optimize existing molecules under a similarity constraint. These are computational property-optimization workflows; the supplied evidence does not establish experimental activity or suitability for a particular discovery program.
The examples are tied to the supplied datasets, model configurations and recorded dependency environment. The README includes author-reported experimental results, but these should not be treated as guaranteed outcomes for other settings. Validity correction is an explicit generation option, so evaluations should distinguish corrected outputs from uncorrected sampling. Compatibility with current software environments and behavior on other datasets remain unestablished by the supplied excerpt.
Key Features
- Preprocesses SMILES strings into molecular graphs for the documented QM9 and zinc250k workflows.
- Provides dataset-specific model-training examples and a link to trained model files.
- Supports molecular reconstruction and random latent-space generation with configurable temperature, sample count and validity correction.
- Includes two-molecule and grid-based latent-space interpolation workflows with molecular visualization.
- Trains latent-space MLP regressors for QED or plogp and uses them to produce ranked optimized molecular candidates.
- Supports constrained property optimization with a configurable molecular similarity cutoff.
Use Cases
- Intended evaluation: assess reconstruction and generation on QM9 or zinc250k, recording validity, novelty and uniqueness separately for corrected and uncorrected outputs.
- Intended evaluation: explore structural changes along two-point or grid-based latent-space interpolations using the documented molecular visualizations.
- Intended evaluation: generate QED-ranked candidates or examine plogp optimization under different similarity constraints, without treating computational scores as experimental validation.
How to Use
- Read the repository README and the MoFlow paper to select a documented dataset and experiment: reconstruction, random generation, interpolation or optimization.
- Follow the README’s environment and dependency instructions. Treat its recorded package versions as historical setup information, not proof of compatibility with a current environment.
- Prepare QM9 or zinc250k through the documented
data_preprocess.pyworkflow, which converts SMILES strings to molecular graphs. Keep the dataset choice consistent across preprocessing, training and generation. - Train with the corresponding
train_model.pyexample or obtain files from the linked trained-model folder. Match the snapshot andmoflow-params.jsonconfiguration to the selected model. - Use the relevant
generate.pyworkflow. For an intended evaluation, record sampling settings, whether validity correction is enabled, and the resulting logs or molecular figures. - For optimization, follow
optimize_property.pyto train or select the property regressor. Check the property-model path and similarity settings before interpreting ranked candidates or constrained-optimization outputs.