Overview
Matchms is a Python library for processing tandem mass spectrometry (MS/MS) data and comparing reference and query spectra. It provides the preparation and scoring stages of a spectral-analysis workflow rather than a standalone predictive model. The documented inputs include mzML, mzXML, msp, MGF, JSON and spectra retrieved through metabolomics-USI. Outputs include processed spectra, pairwise similarity scores and ranked matches for individual queries; the pipeline examples also show writing cleaned query and reference spectra to MGF files.
Workflows can combine metadata standardization, validation, intensity normalization and peak filtering before similarity calculation. The Pipeline class supports creating, importing, exporting and modifying YAML workflow definitions, while individual importing, filtering and scoring functions can be used separately when custom processing is needed. Documented scoring options include cosine, modified cosine and neutral-loss comparisons, alongside molecular-fingerprint and metadata-based measures. Sparse score handling stores computed, non-null values, and staged scoring can use an initial measure to narrow subsequent comparisons.
Matchms also supports custom similarity measures. Spec2Vec and MS2DeepScore are separate packages that follow its scoring interface, not built-in predictive models supplied by Matchms itself. The README announces forthcoming changes for matchms 1.0 and labels much of its guidance as the former README, so users should check the documentation against their selected package version. The supplied package configuration restricts Python to versions from 3.10 through 3.14; the README’s suggestion that higher versions might work is not a compatibility guarantee.
Key Features
- Imports spectra from mzML, mzXML, msp, MGF and JSON, with metabolomics-USI retrieval also documented.
- Provides metadata cleaning and validation, intensity normalization and peak-selection filters, including removal of peaks around precursor m/z.
- Defines reusable processing and scoring workflows through the `Pipeline` class and importable, exportable YAML configurations.
- Offers cosine, modified cosine, neutral-loss, molecular-fingerprint and metadata-based spectral comparison methods.
- Stores computed, non-null scores sparsely and supports staged comparisons using initial pre-selection measures.
- Retrieves sorted matches for a query through `scores_by_query` and accepts custom or separately installed similarity measures.
Use Cases
- Suggested evaluation: build a repeatable MS/MS library-cleaning workflow that standardizes metadata and exports cleaned reference spectra.
- Suggested evaluation: compare query spectra with a reference library and inspect ranked cosine or modified-cosine matches and matching-peak counts.
- Suggested evaluation: assess how peak filtering and normalization change spectral similarity results on representative spectra.
- Suggested evaluation: compare built-in scoring methods with separately installed Spec2Vec or MS2DeepScore in a shared processing workflow.
How to Use
- Read the repository guidance and documentation before choosing a package version. Account for the README’s notice about upcoming matchms 1.0 changes rather than assuming development examples match your installation.
- Follow the documented installation procedure in a new environment to reduce dependency clashes. Check the selected package’s Python requirements; the supplied configuration specifies Python 3.10–3.14.
- Start with the supplied pesticides.mgf example, then consult the importing API for your own spectral format.
- Select metadata and peak-processing steps using the filtering API. Save a YAML workflow with
Pipeline, or assemble individual functions for custom processing. - Choose a measure from the similarity API, calculate reference/query scores and inspect ranked matches. As an intended evaluation, compare filtering choices and tolerances on representative inputs before scaling up.