Overview
SELFIES (Self-Referencing Embedded Strings) represents molecular graphs as sequences of symbols. Its Python library is intended to support molecular machine learning, particularly generative workflows that need syntactically and semantically valid graph representations. It supplies representation and conversion functions rather than a standalone trained generative model; the repository also links to a variational autoencoder example.
The main conversion workflow accepts a SMILES string through selfies.encoder and produces a SELFIES string. selfies.decoder translates SELFIES back into SMILES. Supporting functions count and split symbols, derive an alphabet from a collection of SELFIES strings, and convert strings to integer-label or one-hot encodings and back. The documented padding symbol, [nop], is skipped during decoding, allowing padded sequences to be used without adding molecular content.
Semantic constraints determine how symbols are interpreted. Users can change these constraints, including selecting a hypervalent configuration. The README demonstrates that the same encoded input can yield different decoded structures under different constraint settings. Translation attribution is also available, linking output tokens to the input tokens responsible for them in both encoding and decoding.
The README illustrates random molecule generation using a robust symbol alphabet, while noting that such outputs can require further filtering. Representation validity should therefore not be equated with chemical stability or suitability for an application. Its hypervalence example specifically warns that the stability and reasonableness of many such molecules remain uncertain. Suggested downstream evaluations should examine decoded structures under explicitly recorded constraints rather than treating conversion alone as scientific validation.
Key Features
- Bidirectional translation between SMILES and SELFIES through `selfies.encoder` and `selfies.decoder`.
- Configurable semantic constraints through `selfies.set_semantic_constraints`, with a documented hypervalent configuration.
- Symbol counting, tokenization, and dataset-derived alphabet construction using `selfies.len_selfies`, `selfies.split_selfies`, and `selfies.get_alphabet_from_selfies`.
- Conversion between SELFIES strings and integer-label or one-hot encodings, including padding with the decoder-ignored `[nop]` symbol.
- A robust symbol alphabet accessible through `selfies.get_semantic_robust_alphabet` for the documented random-generation workflow.
- Encoding and decoding attribution that connects output tokens to contributing input tokens, with underlying graph attribution available through `get_attribution`.
Use Cases
- Suggested evaluation: prepare a molecular dataset as padded label or one-hot sequences for a generative model, using a vocabulary derived from the dataset.
- Suggested evaluation: encode representative SMILES and decode them under recorded semantic constraints to assess representation behavior before integrating SELFIES into a cheminformatics pipeline.
- Suggested evaluation: generate candidate strings from the robust alphabet, decode them, and apply application-specific filtering rather than assuming that valid graphs are useful molecules.
- Suggested evaluation: inspect translation attribution to understand how branch, ring, and atom symbols contribute to decoded SMILES.
How to Use
- Read the repository README and its linked code-paper tutorial to understand conversion functions, encodings, and semantic constraints.
- Follow the README’s pip installation instructions for
selfies. Record the installed version and consult the CHANGELOG before changing releases. - For an initial evaluation, select representative SMILES inputs and use
selfies.encoderfollowed byselfies.decoder. Inspect the decoded structures and account for the documentedEncoderErrorandDecoderErrorexceptions. - Build a vocabulary with
selfies.get_alphabet_from_selfies, determine sequence lengths, and useselfies.selfies_to_encodingfor label or one-hot inputs. Include[nop]when padding is needed, and retain the vocabulary mapping for reverse conversion. - Record the semantic constraint configuration, especially when exploring hypervalence. Use attribution to investigate translations, then examine the variational autoencoder example if evaluating model integration. Assess generated structures separately for the intended chemistry task.