Skip to content

JT-VAE

JT-VAE is the official Junction Tree Variational Autoencoder implementation for molecular graph generation, with VAE training code and scripts for Bayesian optimization and joint property-predictor training.

Catalog updated ·

Overview

JT-VAE provides the official implementation of the Junction Tree Variational Autoencoder for Molecular Graph Generation, linked to the paper at https://arxiv.org/abs/1802.04364. Its workflow role is molecular graph generation and associated model-training experiments. The README identifies molecular graphs as the generation target, but the source excerpts do not specify dataset preparation, accepted input formats, output serialization or checkpoint availability.

The repository separates model implementation from training and experiment scripts. The README directs users to fast_jtnn/ for the accelerated implementation and fast_molvae/ for VAE training. It also identifies directories supporting experiments from the original ICML paper: bo/ for Bayesian optimization, molvae/ for VAE-only training, molopt/ for joint training of the VAE and property predictors, and jtnn/ for the model formulation. These are documented workflow components, not evidence of reproduced results.

The listed requirements are Linux, Python 2.7, RDKit version 2017.09 or later, and PyTorch version 0.2 or later; the authors report testing only on Ubuntu. These historical requirements warrant environment planning before evaluation. The README also recommends the separate hgraph2graph repository for architecture improvements. Its described ChEMBL-pretrained model and property-guided generation scripts should not be attributed to this JT-VAE repository. Although the README calls the newer implementation accelerated, the supplied evidence provides no numerical speed comparison.

Key Features

  • Official implementation of the Junction Tree Variational Autoencoder for molecular graph generation, with a linked research paper.
  • Accelerated model implementation in `fast_jtnn/`, paired with VAE training code in `fast_molvae/`.
  • VAE-only training scripts in `molvae/` for experiments associated with the original ICML paper.
  • Joint VAE and property-predictor training scripts in `molopt/`.
  • Bayesian optimization experiment scripts in `bo/`, with a directory-specific README identified by the main documentation.

Use Cases

  • Intended evaluation: investigate JT-VAE as a molecular graph generation model using the documented implementation and training workflow.
  • Intended evaluation: examine joint representation learning and molecular property prediction through the `molopt/` training scripts.
  • Intended evaluation: study the repository's Bayesian optimization experiments as a starting point for a molecular optimization research workflow.

How to Use

  1. Read the official README and linked paper to establish the model's purpose and distinguish implementation instructions from research claims.
  2. Assess the listed environment requirements: Linux, Python 2.7, RDKit >= 2017.09 and PyTorch >= 0.2. Consult the linked RDKit installation documentation; the README recommends conda and reports Ubuntu-only testing.
  3. In the repository, inspect fast_jtnn/ and read fast_molvae/README.md before attempting accelerated VAE training. Confirm input preparation and outputs there rather than assuming formats.
  4. Select the experiment directory appropriate to your evaluation: bo/, molvae/ or molopt/. Read its README for the actual procedure; commands are not supplied in the excerpts.
  5. Review the repository code licence. If architecture improvements are relevant, separately inspect hgraph2graph without treating its capabilities as JT-VAE features.

Related resources

AIDDISON Explorer is a hosted drug-discovery platform that generates and ranks molecular candidates against target profiles, design constraints, predicted properties and synthetic feasibility.

Molecular Generation · Drug Discovery

Chemistry42

Platforms

Chemistry42 is a small-molecule discovery platform combining generative design, retrosynthesis, ADMET and selectivity prediction, and physics-based prioritization for hit identification and lead optimization.

Molecular Generation · Drug Discovery

Official research code for E(3)-equivariant diffusion models that generate 3D molecules, with QM9 and GEOM-Drugs training workflows, sample analysis, and property-conditioned generation.

Open sourcePython

Molecular Generation

FEgrow

Open Source

FEgrow supports interactive ligand-series construction for free-energy preparation, with documented active-learning examples and Dask-based acceleration for molecular design workflows.

Open source

Molecular Generation · Drug Discovery

GeoDiff

Model

GeoDiff is a geometric diffusion model for molecular conformation generation, with official code for GEOM-based training, checkpoint sampling, and conformation and property evaluation.

Open sourcePython

Molecular Generation · Computational Chemistry

GEOM

Dataset

GEOM provides 37 million energy- and statistical-weight-annotated molecular conformations for over 450,000 molecules, with MessagePack data, RDKit objects, and loading and analysis tutorials.

Python

Molecular Generation · Computational Chemistry

Related guides

Workflows

AI for Drug Discovery: Agents, MCP Servers and Platforms

Choose tools by workflow stage: target evidence, protein structures, known binding sites, molecular generation, property analysis and screening. Compare research agents, MCP integrations, hosted APIs and platforms, with a proposed end-to-end example and practical evaluation gates.

Workflows

Reproducibility for Chemistry AI Workflows

Build a compact evaluation record that connects scientific questions to data provenance, transformations, model configurations, outputs and failures, with practical planning examples for chemistry AI.