Overview
Spec2Vec provides a learned similarity measure for tandem mass spectrometry spectra. Inspired by Word2Vec, it learns relationships among mass fragments and neutral losses from their occurrence across a spectral collection, then uses spectral embeddings for comparison. The repository supplies a Python package for training and scoring workflows rather than a single fixed predictive model or a hosted spectral database.
The documented workflow starts with reference spectra loaded from MGF files using matchms. After filtering, SpectrumDocument converts peaks into tokens based on their m/z values, with configurable decimal precision. train_new_word2vec_model trains a model from these documents and can save it for later use. For comparison, a trained model is loaded through gensim and passed to Spec2Vec. Integration with matchms supports scoring combinations of reference and query spectra and retrieving the highest-scoring matches for a selected query. Outputs therefore include a saved Word2Vec model and spectral similarity scores with ranked reference matches.
Model suitability depends on the training collection: the README recommends a large, representative dataset for meaningful representations. Spectra from a different collection may contain peak tokens absent from the model vocabulary, so the scoring interface exposes allowed_missing_percentage. The example also configures intensity weighting. These controls are relevant when evaluating a new dataset, but the source excerpts do not establish benchmark performance or identification accuracy. matchms handles the illustrated import and filtering stages; the README distinguishes RDKit requirements for certain matchms filters from Spec2Vec functionality itself.
Key Features
- Learns Word2Vec-based embeddings from relationships among mass fragments and neutral losses in MS/MS spectra.
- Converts spectra into m/z-based token documents with `SpectrumDocument`, including configurable decimal precision.
- Trains and saves Word2Vec models through `train_new_word2vec_model`, with default training parameters that can be overridden.
- Calculates model-based spectral similarity with configurable `intensity_weighting_power` and `allowed_missing_percentage`.
- Integrates with matchms to score reference–query combinations and retrieve sorted matches for individual query spectra.
Use Cases
- Intended evaluation: rank reference-library spectra for experimental MS/MS queries and assess the relevance of the highest-scoring matches.
- Intended evaluation: train a similarity model on a representative in-house spectral collection and assess its suitability for related query data.
- Intended evaluation: examine how unseen peak tokens and the missing-token allowance affect comparisons when training and query collections differ.
How to Use
- Start with the user documentation and the repository README. Choose the documented Conda or pip installation route; note the separate RDKit consideration if using matchms filters that require it.
- Prepare representative reference spectra and query spectra. Follow the README’s MGF import example with matchms, apply the illustrated filtering pipeline, and exclude spectra rejected by that pipeline.
- Convert cleaned reference spectra into
SpectrumDocumentobjects. Select the m/z decimal precision deliberately, since this determines the peak-token representation used for training. - Train and save a model using
train_new_word2vec_model, or load a model previously trained for your workflow through gensim. The README recommends a large, representative training collection. - Configure
Spec2Vecwith the model, intensity weighting, and an explicit missing-token allowance. Follow the documented matchms scoring example to compare references with queries and retrieve sorted matches. As an intended evaluation, inspect results on representative queries before adopting the scores in a scientific workflow.