Overview
CGCNN is a software implementation of Crystal Graph Convolutional Neural Networks for predicting material properties from crystal structures. It supports two main workflows: training a model on a user-supplied structure–property dataset and applying a pre-trained model to new crystals. The repository documents both regression and classification, with examples for formation energy per atom and metal-versus-semiconductor classification. It serves as a predictive modeling component rather than a database of materials.
Both workflows require a dataset directory containing crystal structures as ID.cif files, a two-column id_prop.csv associating crystal identifiers with target values, and an atom_init.json containing element initialization vectors. For inference, actual target measurements are unnecessary, but the target column must still contain placeholder numbers. The supplied sample datasets illustrate regression and classification inputs; advanced users can also implement a custom PyTorch dataset rather than use the documented CIFData interface.
Training supports separate training, validation, and test partitions, specified either by counts or ratios. Its outputs include a best-validation model, a final-epoch checkpoint, and a CSV of test identifiers, targets, and predictions. Pre-trained inference also produces test_results.csv; classification predictions are probabilities for class 1. The README lists PyTorch, scikit-learn, and pymatgen as dependencies. The source excerpts do not establish predictive accuracy on a prospective dataset or suitability across all crystal families, so target-specific evaluation remains necessary before using predictions to guide materials selection.
Key Features
- Trains CGCNN models on custom datasets pairing CIF crystal structures with material-property targets.
- Supports regression and classification, with separate sample datasets for each task.
- Applies pre-trained checkpoints to new structures, including documented formation-energy and metal-versus-semiconductor examples.
- Allows training, validation, and test partitions to be configured by counts or ratios; these two configuration methods cannot be combined.
- Writes best-validation and final-epoch model checkpoints, plus CSV results containing crystal identifiers, target values, and predictions.
- Provides the `CIFData` dataset interface and describes custom PyTorch dataset classes as an alternative input route.
Use Cases
- Suggested evaluation: train a property regressor on a labeled crystal collection and assess held-out predictions before using it for candidate prioritization.
- Suggested evaluation: apply the supplied formation-energy checkpoint to representative CIF structures and compare predictions with available reference values.
- Suggested evaluation: examine metal-versus-semiconductor classification probabilities on a labeled crystal set to assess suitability for an electronic-materials screening workflow.
How to Use
- Read the official README and choose custom training or pre-trained inference. Review the documented PyTorch, scikit-learn, and pymatgen prerequisites before preparing an environment.
- Assemble a dataset directory with one
ID.cifper crystal, matching identifiers inid_prop.csv, and anatom_init.jsonelement-vector file. Use the documented sample datasets as formatting references. Inference still requires a numeric target column, even when measurements are unavailable. - For training, select regression or classification and define training, validation, and test partitions. Follow the README's count-based or ratio-based configuration, without mixing the two methods.
- For inference, select a pre-trained checkpoint appropriate to the target property and follow the README's
predict.pyworkflow. The documented examples cover formation energy per atom and metal-versus-semiconductor classification. - Inspect
test_results.csvand, after training, the saved checkpoints. Treat inference placeholders as non-measurements and classification outputs as class-1 probabilities. As a suggested evaluation step, compare predictions with independent reference labels before downstream use; consult the framework paper for methodological context.