Overview
Molecular Sets (MOSES) is a benchmarking platform for molecular generation research, combining reference data, model implementations and evaluation tools. It supports a shared workflow for training generators, sampling molecular structures and comparing generated collections with held-out data. It is therefore more than a standalone dataset or a single predictive model.
The benchmark contains 1,936,962 molecular structures refined from the ZINC Clean Leads collection. Its preparation applies molecular-weight, rotatable-bond and XlogP limits, excludes specified atom and ring configurations, and uses medicinal chemistry and PAINS filters. The resulting data are divided into approximately 1.6 million training molecules and two evaluation sets of about 176,000 molecules each. The scaffold test set holds Bemis-Murcko scaffolds absent from both the training and ordinary test sets, providing a separate reference for assessing generation of unseen scaffolds.
For custom models, the documented interface supplies dataset splits and accepts generated SMILES for metric calculation. Evaluation covers validity, uniqueness, novelty, nearest-neighbor similarity, internal diversity, fragment and scaffold comparisons, and Fréchet ChemNet Distance. The baseline workflow also supports dataset splitting, training, generation and evaluation, with results written to metrics.csv.
Interpretation should account for the benchmark’s filtered chemical scope: comparisons concern this reference collection, not unrestricted chemical space. The README recommends at least 30,000 generated molecules and repeated experiments to estimate metric variance. These outputs characterize generated collections; they should not be treated as experimental confirmation of biological activity or synthetic feasibility.
Key Features
- Provides a filtered ZINC-derived benchmark with training, test and scaffold-held-out test splits accessible through `moses.get_dataset`.
- Documents generation baselines including CharRNN, VAE, AAE and LatentGAN, and links to a JTN-VAE implementation.
- Calculates validity, uniqueness, novelty, SNN, IntDiv, fragment and scaffold metrics, and Fréchet ChemNet Distance from generated SMILES.
- Supplies `CharVocab` and `StringDataset` helpers for preparing string-based molecular data for a torch DataLoader.
- Includes an end-to-end baseline workflow that splits data, trains models, generates molecules and saves evaluation results to `metrics.csv`.
- Documents separate training, sampling and evaluation interfaces, plus a workflow for evaluating models across multiple seeds.
Use Cases
- Suggested evaluation: compare a new molecular generator with documented baselines using the same reference splits and metric suite.
- Suggested evaluation: examine differences between ordinary test-set comparisons and scaffold-held-out comparisons when assessing structural generalization.
- Suggested evaluation: measure validity, uniqueness and novelty alongside distributional similarity to identify trade-offs in generated molecular collections.
- Suggested evaluation: repeat baseline or custom-model experiments to estimate variation in evaluation metrics across runs.
How to Use
- Read the official README to choose between custom-model benchmarking and baseline reproduction. Consult the paper for the benchmark’s research context.
- Follow the README’s installation instructions, including RDKit setup. The package is named
molsets; LatentGAN has additional documented dependencies. Select the installation route appropriate to your environment rather than assuming current compatibility. - Obtain the
train,testandtest_scaffoldssplits throughmoses.get_dataset, or inspect the linked benchmark dataset. Keep the held-out references separate from training inputs. - Train your generator or follow a documented baseline workflow, then collect generated SMILES. The README recommends sampling at least 30,000 molecules for metric calculation.
- Evaluate samples with
moses.get_all_metricsor the documented evaluation interface. Retain generated samples and metric outputs, repeat experiments to estimate variance, and interpret results within the benchmark’s filtered chemical domain.