Skip to content

ChemBERTa

ChemBERTa provides BERT-like models for chemical SMILES, with RoBERTa masked-language modelling checkpoints and notebooks for pre-training, fine-tuning and molecular property prediction research.

Catalog updated ·

Overview

ChemBERTa is a collection of BERT-like models applied to chemical SMILES strings. The repository positions these models for chemical modelling, drug-design research and molecular property prediction. Its documented workflow centres on notebooks for pre-training and fine-tuning, rather than a standalone chemistry application. The README identifies RoBERTa, a BERT variant, as the architecture used in the notebooks and masked-language modelling as their training task.

SMILES text is the model input. The supplied example loads a matching tokenizer and model from the checkpoint seyonec/ChemBERTa-zinc-base-v1, then constructs a Hugging Face fill-mask pipeline for masked-token prediction. The README also links to pretrained weights associated with ZINC 100k, ZINC 250k, PubChem 100k, PubChem 250k, PubChem 1M and PubChem 10M. These checkpoints provide starting points for investigating chemical language modelling and downstream adaptation; the repository's notebook workflow includes evaluation and fine-tuning.

For property-prediction work, the linked DeepChem transfer-learning tutorial is a practical entry point, while the ChemBERTa and ChemBERTa-2 papers provide research references. Masked-token predictions should be distinguished from task-specific molecular property outputs, which require an appropriate downstream workflow. The README mixes notebook descriptions with roadmap and checklist statements, including work described as in progress. Those statements do not establish that every proposed implementation or visualization component is ready to use, and the source excerpts do not establish current dependency compatibility or predictive performance.

Key Features

  • RoBERTa-based chemical language modelling using SMILES strings and a masked-language modelling training objective.
  • Pretrained weights linked for six ZINC and PubChem dataset subsets, ranging from 100k to 10M as listed in the README.
  • Notebooks documenting pre-training, fine-tuning and evaluation workflows for ChemBERTa models.
  • A tokenizer-and-model loading example using `seyonec/ChemBERTa-zinc-base-v1` with a Hugging Face `fill-mask` pipeline.
  • A linked DeepChem tutorial focused on transfer learning with ChemBERTa transformers.

Use Cases

  • Suggested evaluation: investigate masked-token predictions on representative SMILES inputs using the documented tokenizer and checkpoint pairing.
  • Suggested evaluation: assess ChemBERTa fine-tuning for a molecular property dataset through the linked transfer-learning workflow.
  • Suggested evaluation: compare the listed ZINC- and PubChem-pretrained checkpoints under a consistent downstream evaluation protocol.

How to Use

  1. Read the repository README to distinguish the documented notebook workflow from roadmap items. Decide whether your goal is masked-language modelling or downstream property prediction.
  2. Consult the DeepChem transfer-learning tutorial for the linked learning workflow. Check its actual requirements before preparing an environment; the source excerpts do not specify installation steps.
  3. Browse the linked Hugging Face models and select a checkpoint relevant to your evaluation. Inspect checkpoint-specific documentation and terms separately from the repository code licence.
  4. For masked-token work, follow the README's model-and-tokenizer pairing for seyonec/ChemBERTa-zinc-base-v1 and its fill-mask pipeline example. Evaluate behaviour on representative SMILES rather than assuming property-prediction output.
  5. For downstream experiments, use the tutorial to guide fine-tuning and evaluation. Consult the ChemBERTa-2 paper for the research reference requested by the README.

Related resources

ChemML

Open Source

ChemML is a Python suite for chemical and materials data analysis, mining, and modeling, with a modular design and documented work on graph neural networks, AutoML, and explainability.

Open sourcePython

Molecular Property Prediction · Materials Discovery

Chemprop is a PyTorch-based toolkit for training and evaluating message-passing neural networks for molecular property prediction, with CLI workflows, Python modules and task-specific notebooks.

Open sourcePython

Molecular Property Prediction

DeepMol

Open Source

DeepMol is a Python toolkit for molecular machine learning, connecting compound standardization and featurization with model training, evaluation, interpretation and reusable pipelines.

Open sourcePython

Molecular Property Prediction

DGL-LifeSci

Open Source

DGL-LifeSci is a DGL/PyTorch toolkit providing graph learning components and examples for chemistry and biology.

Open sourcePython

Molecular Property Prediction

DimeNet

Model

DimeNet provides reference implementations of DimeNet and DimeNet++ for directional message passing on molecular graphs, with training notebooks, test-set prediction workflows and pretrained models.

Python

Molecular Property Prediction · Quantum Chemistry

DrugAgent is a research multi-agent framework that combines LLM planning with domain-guided code generation to build and evaluate machine-learning workflows for drug discovery.

Molecular Property Prediction · Drug Discovery

Related guides