Skip to content

Spec2Vec

Florian Huber / Netherlands eScience Center & Düsseldorf University of Applied Sciences

Spec2Vec is a Python package for MS/MS spectral similarity scoring, using Word2Vec embeddings learned from mass fragments and neutral losses to compare query and reference spectra.

Catalog updated ·

Overview

Spec2Vec provides a learned similarity measure for tandem mass spectrometry spectra. Inspired by Word2Vec, it learns relationships among mass fragments and neutral losses from their occurrence across a spectral collection, then uses spectral embeddings for comparison. The repository supplies a Python package for training and scoring workflows rather than a single fixed predictive model or a hosted spectral database.

The documented workflow starts with reference spectra loaded from MGF files using matchms. After filtering, SpectrumDocument converts peaks into tokens based on their m/z values, with configurable decimal precision. train_new_word2vec_model trains a model from these documents and can save it for later use. For comparison, a trained model is loaded through gensim and passed to Spec2Vec. Integration with matchms supports scoring combinations of reference and query spectra and retrieving the highest-scoring matches for a selected query. Outputs therefore include a saved Word2Vec model and spectral similarity scores with ranked reference matches.

Model suitability depends on the training collection: the README recommends a large, representative dataset for meaningful representations. Spectra from a different collection may contain peak tokens absent from the model vocabulary, so the scoring interface exposes allowed_missing_percentage. The example also configures intensity weighting. These controls are relevant when evaluating a new dataset, but the source excerpts do not establish benchmark performance or identification accuracy. matchms handles the illustrated import and filtering stages; the README distinguishes RDKit requirements for certain matchms filters from Spec2Vec functionality itself.

Key Features

  • Learns Word2Vec-based embeddings from relationships among mass fragments and neutral losses in MS/MS spectra.
  • Converts spectra into m/z-based token documents with `SpectrumDocument`, including configurable decimal precision.
  • Trains and saves Word2Vec models through `train_new_word2vec_model`, with default training parameters that can be overridden.
  • Calculates model-based spectral similarity with configurable `intensity_weighting_power` and `allowed_missing_percentage`.
  • Integrates with matchms to score reference–query combinations and retrieve sorted matches for individual query spectra.

Use Cases

  • Intended evaluation: rank reference-library spectra for experimental MS/MS queries and assess the relevance of the highest-scoring matches.
  • Intended evaluation: train a similarity model on a representative in-house spectral collection and assess its suitability for related query data.
  • Intended evaluation: examine how unseen peak tokens and the missing-token allowance affect comparisons when training and query collections differ.

How to Use

  1. Start with the user documentation and the repository README. Choose the documented Conda or pip installation route; note the separate RDKit consideration if using matchms filters that require it.
  2. Prepare representative reference spectra and query spectra. Follow the README’s MGF import example with matchms, apply the illustrated filtering pipeline, and exclude spectra rejected by that pipeline.
  3. Convert cleaned reference spectra into SpectrumDocument objects. Select the m/z decimal precision deliberately, since this determines the peak-token representation used for training.
  4. Train and save a model using train_new_word2vec_model, or load a model previously trained for your workflow through gensim. The README recommends a large, representative training collection.
  5. Configure Spec2Vec with the model, intensity weighting, and an explicit missing-token allowance. Follow the documented matchms scoring example to compare references with queries and retrieve sorted matches. As an intended evaluation, inspect results on representative queries before adopting the scores in a scientific workflow.

Related resources

Matchms

Open Source

Matchms is a Python toolkit for importing, cleaning and comparing tandem mass spectra, with configurable processing pipelines, multiple similarity measures and sparse score storage.

Open sourcePython

Spectroscopy

MS2DeepScore predicts molecular structural similarity from pairs of tandem mass spectra using a Siamese neural network, with tools for spectral embeddings, inference and custom model training.

Open sourcePython

Spectroscopy

A K-Dense Skill with instructions and a Python helper for calibrated 1D NMR FID processing, producing phased spectra, positive peak candidates, signed integrals, and reproducible processing reports.

Open sourcePython

Spectroscopy

Related guides

Skills

AI Skills for Chemistry: What They Are and How They Work

Understand chemistry AI Skills as reusable procedure modules: what SKILL.md contains, how hosts load instructions, how Skills differ from MCP and agents, and how to select, combine, and evaluate them without confusing guidance with scientific validation.