Skip to content

DScribe

DScribe is a Python library that converts atomic structures into numerical descriptors for machine learning, visualization and similarity analysis, with batch processing and atomic-position derivatives.

Catalog updated ·

Overview

DScribe provides feature representations for atomistic workflows in materials science. It transforms atomic structures into fixed-size numerical fingerprints that can serve as inputs to machine learning, visualization or similarity analysis. Its role is descriptor generation rather than a supplied property-prediction model: researchers can use the resulting features in a separate downstream analysis or learning workflow.

The README demonstrates inputs constructed with ASE, using H2O, NO2 and CO2 as examples. Descriptor objects are configured before structures are passed to their creation methods. The CoulombMatrix example specifies a maximum atom count and a sorting option, while the SOAP example defines chemical species, a cutoff and expansion parameters. Outputs can be NumPy arrays or sparse arrays. SOAP calculations can target selected atomic centers, including oxygen atoms selected across multiple structures, and batches can be processed across multiple processes.

The documented descriptor collection includes Coulomb, Sine and Ewald matrices; Atom-centered Symmetry Functions (ACSF); Smooth Overlap of Atomic Positions (SOAP); Many-body Tensor Representation (MBTR); Local Many-body Tensor Representation (LMBTR); and the Valle-Oganov descriptor. The README lists atomic-position derivative support for all eight and demonstrates returning SOAP derivatives together with descriptors. The source excerpts do not establish predictive accuracy, benchmark speed or the best descriptor settings for a particular scientific problem. Descriptor choice and configuration therefore remain matters for intended evaluation on representative structures; the linked documentation provides the route to further tutorials and installation details.

Key Features

  • Generates numerical fingerprints from atomic structures, with NumPy-array or sparse-array outputs.
  • Implements eight descriptor families: Coulomb matrix, Sine matrix, Ewald matrix, ACSF, SOAP, MBTR, LMBTR and Valle-Oganov.
  • Supports SOAP evaluation at selected atomic centers, including center selections supplied for multiple structures.
  • Creates descriptors for individual structures or batches, with multi-process processing controlled through n_jobs.
  • Provides derivatives with respect to atomic positions; the SOAP example returns derivatives and descriptors together.

Use Cases

  • Intended evaluation: compare descriptor families as input features for a separate molecular- or materials-property prediction model.
  • Intended evaluation: generate fingerprints for visualization or similarity analysis of a collection of atomic structures.
  • Intended evaluation: examine SOAP representations at selected element-specific atomic centers across several structures.
  • Intended evaluation: use descriptor derivatives in a downstream workflow that requires sensitivity to changes in atomic positions.

How to Use

  1. Start with the official documentation to identify the descriptor and tutorial relevant to your structures. Treat the README examples as starting points rather than validated settings for your application.
  2. Follow the installation guide, choosing among the pip, conda or source routes documented in the README. The supplied package metadata declares Python >=3.9; this is a requirement, not evidence of tested compatibility.
  3. Inspect the repository example for ASE-based structure preparation. Assemble representative structures and identify any atomic centers you want to describe.
  4. Configure the descriptor. For CoulombMatrix, review the maximum atom count and permutation setting; for SOAP, review species, cutoff and expansion parameters. Generate a single-structure output before moving to batches.
  5. Evaluate batch creation, optional multi-process processing and, where needed, atomic-position derivatives. Check output dimensions and downstream suitability on your own data before using the fingerprints for prediction, visualization or similarity analysis.

Related resources

ChemML

Open Source

ChemML is a Python suite for chemical and materials data analysis, mining, and modeling, with a modular design and documented work on graph neural networks, AutoML, and explainability.

Open sourcePython

Molecular Property Prediction · Materials Discovery

MEGNet

Model

MEGNet is a deprecated TensorFlow graph-network implementation with pretrained molecular and crystal property models, training utilities, and transferable elemental embeddings.

Open sourcePython

Molecular Property Prediction · Materials Discovery

ALIGNN

Model

ALIGNN provides atomistic graph neural networks for materials property prediction, with training workflows, pretrained predictors and ALIGNN-FF force fields for structural optimization.

Open sourcePython

Materials Discovery

Allegro

Model

Allegro implements an E(3)-equivariant interatomic potential as a NequIP extension, with documented GPU acceleration options and a separate plugin for LAMMPS simulations.

Open sourcePython

Computational Chemistry · Materials Discovery

An ASE routing Skill in the computational-chemistry-agent-skills collection that separates workflow preparation from calculator configuration and delegates execution elsewhere.

Computational Chemistry · Materials Discovery

CGCNN

Model

CGCNN implements crystal graph convolutional neural networks for learning material properties from crystal structures, with custom-data training and prediction using pre-trained models.

Open sourcePython

Materials Discovery

Related guides