Overview
DScribe provides feature representations for atomistic workflows in materials science. It transforms atomic structures into fixed-size numerical fingerprints that can serve as inputs to machine learning, visualization or similarity analysis. Its role is descriptor generation rather than a supplied property-prediction model: researchers can use the resulting features in a separate downstream analysis or learning workflow.
The README demonstrates inputs constructed with ASE, using H2O, NO2 and CO2 as examples. Descriptor objects are configured before structures are passed to their creation methods. The CoulombMatrix example specifies a maximum atom count and a sorting option, while the SOAP example defines chemical species, a cutoff and expansion parameters. Outputs can be NumPy arrays or sparse arrays. SOAP calculations can target selected atomic centers, including oxygen atoms selected across multiple structures, and batches can be processed across multiple processes.
The documented descriptor collection includes Coulomb, Sine and Ewald matrices; Atom-centered Symmetry Functions (ACSF); Smooth Overlap of Atomic Positions (SOAP); Many-body Tensor Representation (MBTR); Local Many-body Tensor Representation (LMBTR); and the Valle-Oganov descriptor. The README lists atomic-position derivative support for all eight and demonstrates returning SOAP derivatives together with descriptors. The source excerpts do not establish predictive accuracy, benchmark speed or the best descriptor settings for a particular scientific problem. Descriptor choice and configuration therefore remain matters for intended evaluation on representative structures; the linked documentation provides the route to further tutorials and installation details.
Key Features
- Generates numerical fingerprints from atomic structures, with NumPy-array or sparse-array outputs.
- Implements eight descriptor families: Coulomb matrix, Sine matrix, Ewald matrix, ACSF, SOAP, MBTR, LMBTR and Valle-Oganov.
- Supports SOAP evaluation at selected atomic centers, including center selections supplied for multiple structures.
- Creates descriptors for individual structures or batches, with multi-process processing controlled through n_jobs.
- Provides derivatives with respect to atomic positions; the SOAP example returns derivatives and descriptors together.
Use Cases
- Intended evaluation: compare descriptor families as input features for a separate molecular- or materials-property prediction model.
- Intended evaluation: generate fingerprints for visualization or similarity analysis of a collection of atomic structures.
- Intended evaluation: examine SOAP representations at selected element-specific atomic centers across several structures.
- Intended evaluation: use descriptor derivatives in a downstream workflow that requires sensitivity to changes in atomic positions.
How to Use
- Start with the official documentation to identify the descriptor and tutorial relevant to your structures. Treat the README examples as starting points rather than validated settings for your application.
- Follow the installation guide, choosing among the pip, conda or source routes documented in the README. The supplied package metadata declares Python >=3.9; this is a requirement, not evidence of tested compatibility.
- Inspect the repository example for ASE-based structure preparation. Assemble representative structures and identify any atomic centers you want to describe.
- Configure the descriptor. For CoulombMatrix, review the maximum atom count and permutation setting; for SOAP, review species, cutoff and expansion parameters. Generate a single-structure output before moving to batches.
- Evaluate batch creation, optional multi-process processing and, where needed, atomic-position derivatives. Check output dimensions and downstream suitability on your own data before using the fingerprints for prediction, visualization or similarity analysis.