Skip to content

SELFIES

Mario Krenn, Alston Lo, Robert Pollice and many other contributors

SELFIES is a molecular string representation and Python library for translating SMILES, preparing token encodings, and generating molecular graphs under configurable semantic constraints.

Catalog updated ·

Overview

SELFIES (Self-Referencing Embedded Strings) represents molecular graphs as sequences of symbols. Its Python library is intended to support molecular machine learning, particularly generative workflows that need syntactically and semantically valid graph representations. It supplies representation and conversion functions rather than a standalone trained generative model; the repository also links to a variational autoencoder example.

The main conversion workflow accepts a SMILES string through selfies.encoder and produces a SELFIES string. selfies.decoder translates SELFIES back into SMILES. Supporting functions count and split symbols, derive an alphabet from a collection of SELFIES strings, and convert strings to integer-label or one-hot encodings and back. The documented padding symbol, [nop], is skipped during decoding, allowing padded sequences to be used without adding molecular content.

Semantic constraints determine how symbols are interpreted. Users can change these constraints, including selecting a hypervalent configuration. The README demonstrates that the same encoded input can yield different decoded structures under different constraint settings. Translation attribution is also available, linking output tokens to the input tokens responsible for them in both encoding and decoding.

The README illustrates random molecule generation using a robust symbol alphabet, while noting that such outputs can require further filtering. Representation validity should therefore not be equated with chemical stability or suitability for an application. Its hypervalence example specifically warns that the stability and reasonableness of many such molecules remain uncertain. Suggested downstream evaluations should examine decoded structures under explicitly recorded constraints rather than treating conversion alone as scientific validation.

Key Features

  • Bidirectional translation between SMILES and SELFIES through `selfies.encoder` and `selfies.decoder`.
  • Configurable semantic constraints through `selfies.set_semantic_constraints`, with a documented hypervalent configuration.
  • Symbol counting, tokenization, and dataset-derived alphabet construction using `selfies.len_selfies`, `selfies.split_selfies`, and `selfies.get_alphabet_from_selfies`.
  • Conversion between SELFIES strings and integer-label or one-hot encodings, including padding with the decoder-ignored `[nop]` symbol.
  • A robust symbol alphabet accessible through `selfies.get_semantic_robust_alphabet` for the documented random-generation workflow.
  • Encoding and decoding attribution that connects output tokens to contributing input tokens, with underlying graph attribution available through `get_attribution`.

Use Cases

  • Suggested evaluation: prepare a molecular dataset as padded label or one-hot sequences for a generative model, using a vocabulary derived from the dataset.
  • Suggested evaluation: encode representative SMILES and decode them under recorded semantic constraints to assess representation behavior before integrating SELFIES into a cheminformatics pipeline.
  • Suggested evaluation: generate candidate strings from the robust alphabet, decode them, and apply application-specific filtering rather than assuming that valid graphs are useful molecules.
  • Suggested evaluation: inspect translation attribution to understand how branch, ring, and atom symbols contribute to decoded SMILES.

How to Use

  1. Read the repository README and its linked code-paper tutorial to understand conversion functions, encodings, and semantic constraints.
  2. Follow the README’s pip installation instructions for selfies. Record the installed version and consult the CHANGELOG before changing releases.
  3. For an initial evaluation, select representative SMILES inputs and use selfies.encoder followed by selfies.decoder. Inspect the decoded structures and account for the documented EncoderError and DecoderError exceptions.
  4. Build a vocabulary with selfies.get_alphabet_from_selfies, determine sequence lengths, and use selfies.selfies_to_encoding for label or one-hot inputs. Include [nop] when padding is needed, and retain the vocabulary mapping for reverse conversion.
  5. Record the semantic constraint configuration, especially when exploring hypervalence. Use attribution to investigate translations, then examine the variational autoencoder example if evaluating model integration. Assess generated structures separately for the intended chemistry task.

Related resources

AIDDISON Explorer is a hosted drug-discovery platform that generates and ranks molecular candidates against target profiles, design constraints, predicted properties and synthetic feasibility.

Molecular Generation · Drug Discovery

ChemCP is an MCP App that turns SMILES strings into interactive 2D molecular diagrams and computed descriptors using RDKit.js, with an in-chat viewer for hosts that support MCP Apps.

TypeScriptJavaScript

Cheminformatics

ChemCrow is a Python chemistry agent package built with Langchain that connects language-model reasoning to RDKit, paper-qa, chemical databases, and reaction-planning tools.

Open sourcePython

Cheminformatics · Literature & Research

Chemistry42

Platforms

Chemistry42 is a small-molecule discovery platform combining generative design, retrosynthesis, ADMET and selectivity prediction, and physics-based prioritization for hit identification and lead optimization.

Molecular Generation · Drug Discovery

Chemistry Development Kit (CDK) is a Java library for chemical representations, file conversion, structure searching, fingerprints, QSAR descriptors and molecular rendering within other programs.

Open source

Cheminformatics

Datamol

Open Source

Datamol is a Python molecular-processing library built on RDKit, with tools for structure conversion, standardization, fingerprints, conformers, visualization and local or remote file I/O.

Open sourcePython

Cheminformatics

Related guides

Workflows

AI for Drug Discovery: Agents, MCP Servers and Platforms

Choose tools by workflow stage: target evidence, protein structures, known binding sites, molecular generation, property analysis and screening. Compare research agents, MCP integrations, hosted APIs and platforms, with a proposed end-to-end example and practical evaluation gates.

Skills

AI Skills for Chemistry: What They Are and How They Work

Understand chemistry AI Skills as reusable procedure modules: what SKILL.md contains, how hosts load instructions, how Skills differ from MCP and agents, and how to select, combine, and evaluate them without confusing guidance with scientific validation.