Overview
MolGPT is a molecular-generation model project built around a small custom GPT trained through next-token prediction. The README describes training on the MOSES and Guacamol datasets and using the resulting model for both unconditional and conditional molecular generation. Its cited preprint identifies the approach as a transformer-decoder model. In a research workflow, the repository provides an entry point for investigating this generative approach rather than a documented end-to-end molecular design service.
The documented training inputs are processed dataset files in CSV format, which must be placed in the same directory as the code files. The README links these processed datasets, their upstream MOSES and Guacamol repositories, and a separate collection of trained weights. It lists dataset-specific training scripts and two generation scripts. Molecular generation is the stated output task, but the supplied excerpt does not specify the generated representation, output file format, or how conditioning inputs are supplied.
Interpretability is also part of the described work: saliency maps are obtained using the Ecco library. The README states that MolGPT is compared with previous approaches on both datasets, but supplies no numerical results or evaluation protocol. The available evidence therefore supports the training, generation, and saliency-analysis workflow without establishing performance, chemical validity, or practical utility of generated candidates. Installation requirements, hardware needs, and detailed parameter configuration are not documented in the supplied excerpt.
Key Features
- Small custom GPT trained with a next-token prediction objective on MOSES and Guacamol.
- Supports unconditional and conditional molecular generation as described in the README.
- Provides separate training entry points, `train_moses.sh` and `train_guacamol.sh`.
- Lists generation entry points `generate_guacamol_prop.sh` and `generate_moses_prop_scaf.sh`.
- Links processed CSV datasets and a separate collection of trained model weights.
- Describes Ecco-based saliency maps for model interpretability.
Use Cases
- Intended evaluation: investigate unconditional and conditional molecular generation using the linked datasets and trained weights.
- Intended evaluation: study dataset-specific training workflows on MOSES and Guacamol, assessing generated candidates with an independently defined evaluation protocol.
- Intended evaluation: explore Ecco-based saliency analysis to examine model behavior during molecular generation.
How to Use
- Read the official README to distinguish the training and generation entry points. The supplied instructions do not specify dependencies, hardware requirements, or configuration details.
- Obtain the processed CSV datasets from the linked download folder. Consult the upstream Guacamol and MOSES repositories for dataset context.
- For training, place the selected dataset CSV in the same directory as the code files, as instructed. Inspect
train_moses.shortrain_guacamol.shin the repository before choosing the corresponding workflow. - For a weights-based evaluation, consult the trained weights collection. Inspect the generation scripts to determine their expected inputs and weight-loading configuration; these details are absent from the excerpt.
- Select the relevant documented generation entry point and plan independent checks of its outputs. If investigating interpretability, examine the repository’s Ecco-related implementation; the README mentions saliency maps but gives no procedure.