Overview
MS2DeepScore is a Python library and predictive-model workflow for comparing tandem mass spectra. Its Siamese neural network learns to estimate molecular structural similarity, expressed as Tanimoto scores, from spectral pairs. It provides components for preparing data, training models and calculating similarities, making it a scoring component for mass-spectrometry analysis rather than a database of compound identifications.
The README links a pretrained model described as trained on more than 500,000 combined MS/MS spectra from GNPS, Mona, MassBank and MSnLib. That model supports positive and negative ionization modes, including comparisons across modes. The documented workflow uses matchms to load and filter spectra, then returns pairwise predictions as a NumPy similarity matrix. Listed input formats include msp, mzml, mgf, mzxml, json and usi. PyTorch and ONNX inference paths are documented, along with conversion of existing models to ONNX.
Each spectrum can also be represented by an embedding vector for downstream visualization, such as dimensionality reduction with UMAP. An optional embedding evaluator estimates model accuracy for individual spectra; this is a model-derived estimate, not independent validation. Custom training is supported, but the README recommends machine-learning experience, a large and chemically diverse spectral collection, cleaned inputs and inspection of pair sampling. These requirements matter when adapting the model to an in-house dataset: predictions and visualizations should be evaluated on representative spectra before being relied upon in a scientific workflow.
Key Features
- Siamese neural-network scoring that predicts molecular structural similarity as Tanimoto scores from pairs of MS/MS spectra.
- A linked pretrained model supporting positive-mode, negative-mode and cross-ionization-mode spectral comparisons.
- Integration with matchms filtering and scoring pipelines, producing a NumPy array of pairwise similarity predictions.
- Spectrum embedding extraction through both PyTorch and ONNX interfaces for downstream chemical-space visualization.
- Custom model training with configurable settings, additional metadata inputs and optional training of an embedding evaluator.
- ONNX inference support and export of existing PyTorch models to ONNX.
Use Cases
- Suggested evaluation: compare spectra within a metabolomics collection to assess whether predicted structural similarities capture relationships relevant to the study.
- Suggested evaluation: test positive-to-negative ionization-mode comparisons on representative spectral pairs with known molecular structures.
- Suggested evaluation: generate embeddings and inspect UMAP visualizations for exploratory grouping of spectra in chemical space.
- Suggested evaluation: train a model using cleaned in-house spectra combined with public training data, checking pair sampling and prediction quality before downstream use.
How to Use
- Read the repository README for environment preparation and installation options. It lists Python 3.11–3.13; higher versions are not described as systematically tested.
- Follow the linked getting-started tutorial to understand data preparation, scoring and visualization before substituting your own spectra.
- Obtain the pretrained
ms2deepscore_model.ptfrom the model record. Use the README's PyTorch workflow, or its conversion instructions if choosing ONNX inference. - Prepare spectra in a documented format and apply the recommended matchms cleaning workflow. Compute the similarity matrix, then extract embeddings if visualization is needed.
- Evaluate predictions on representative spectral pairs with known structures. For custom training, inspect the pair sampling tutorial and assess dataset diversity before changing training settings.