Skip to content

MolGPT

MolGPT trains a small custom GPT with next-token prediction on MOSES and Guacamol for unconditional and conditional molecular generation, with trained weights and Ecco-based saliency analysis linked.

Catalog updated ·

Overview

MolGPT is a molecular-generation model project built around a small custom GPT trained through next-token prediction. The README describes training on the MOSES and Guacamol datasets and using the resulting model for both unconditional and conditional molecular generation. Its cited preprint identifies the approach as a transformer-decoder model. In a research workflow, the repository provides an entry point for investigating this generative approach rather than a documented end-to-end molecular design service.

The documented training inputs are processed dataset files in CSV format, which must be placed in the same directory as the code files. The README links these processed datasets, their upstream MOSES and Guacamol repositories, and a separate collection of trained weights. It lists dataset-specific training scripts and two generation scripts. Molecular generation is the stated output task, but the supplied excerpt does not specify the generated representation, output file format, or how conditioning inputs are supplied.

Interpretability is also part of the described work: saliency maps are obtained using the Ecco library. The README states that MolGPT is compared with previous approaches on both datasets, but supplies no numerical results or evaluation protocol. The available evidence therefore supports the training, generation, and saliency-analysis workflow without establishing performance, chemical validity, or practical utility of generated candidates. Installation requirements, hardware needs, and detailed parameter configuration are not documented in the supplied excerpt.

Key Features

  • Small custom GPT trained with a next-token prediction objective on MOSES and Guacamol.
  • Supports unconditional and conditional molecular generation as described in the README.
  • Provides separate training entry points, `train_moses.sh` and `train_guacamol.sh`.
  • Lists generation entry points `generate_guacamol_prop.sh` and `generate_moses_prop_scaf.sh`.
  • Links processed CSV datasets and a separate collection of trained model weights.
  • Describes Ecco-based saliency maps for model interpretability.

Use Cases

  • Intended evaluation: investigate unconditional and conditional molecular generation using the linked datasets and trained weights.
  • Intended evaluation: study dataset-specific training workflows on MOSES and Guacamol, assessing generated candidates with an independently defined evaluation protocol.
  • Intended evaluation: explore Ecco-based saliency analysis to examine model behavior during molecular generation.

How to Use

  1. Read the official README to distinguish the training and generation entry points. The supplied instructions do not specify dependencies, hardware requirements, or configuration details.
  2. Obtain the processed CSV datasets from the linked download folder. Consult the upstream Guacamol and MOSES repositories for dataset context.
  3. For training, place the selected dataset CSV in the same directory as the code files, as instructed. Inspect train_moses.sh or train_guacamol.sh in the repository before choosing the corresponding workflow.
  4. For a weights-based evaluation, consult the trained weights collection. Inspect the generation scripts to determine their expected inputs and weight-loading configuration; these details are absent from the excerpt.
  5. Select the relevant documented generation entry point and plan independent checks of its outputs. If investigating interpretability, examine the repository’s Ecco-related implementation; the README mentions saliency maps but gives no procedure.

Related resources

AIDDISON Explorer is a hosted drug-discovery platform that generates and ranks molecular candidates against target profiles, design constraints, predicted properties and synthetic feasibility.

Molecular Generation · Drug Discovery

Chemistry42

Platforms

Chemistry42 is a small-molecule discovery platform combining generative design, retrosynthesis, ADMET and selectivity prediction, and physics-based prioritization for hit identification and lead optimization.

Molecular Generation · Drug Discovery

Official research code for E(3)-equivariant diffusion models that generate 3D molecules, with QM9 and GEOM-Drugs training workflows, sample analysis, and property-conditioned generation.

Open sourcePython

Molecular Generation

FEgrow

Open Source

FEgrow supports interactive ligand-series construction for free-energy preparation, with documented active-learning examples and Dask-based acceleration for molecular design workflows.

Open source

Molecular Generation · Drug Discovery

GeoDiff

Model

GeoDiff is a geometric diffusion model for molecular conformation generation, with official code for GEOM-based training, checkpoint sampling, and conformation and property evaluation.

Open sourcePython

Molecular Generation · Computational Chemistry

GEOM

Dataset

GEOM provides 37 million energy- and statistical-weight-annotated molecular conformations for over 450,000 molecules, with MessagePack data, RDKit objects, and loading and analysis tutorials.

Python

Molecular Generation · Computational Chemistry

Related guides

Workflows

AI for Drug Discovery: Agents, MCP Servers and Platforms

Choose tools by workflow stage: target evidence, protein structures, known binding sites, molecular generation, property analysis and screening. Compare research agents, MCP integrations, hosted APIs and platforms, with a proposed end-to-end example and practical evaluation gates.

Workflows

Reproducibility for Chemistry AI Workflows

Build a compact evaluation record that connects scientific questions to data provenance, transformations, model configurations, outputs and failures, with practical planning examples for chemistry AI.