Overview
GROVER is the PyTorch implementation accompanying Self-Supervised Graph Transformer on Large-Scale Molecular Data. It combines downloadable pretrained models with code for learning molecular representations and adapting them to labelled property-prediction tasks. The README provides GROVER base and large checkpoints, alongside finetuned models for eleven classification and regression datasets. Its workflow spans representation learning, downstream training, prediction, fingerprint extraction and evaluation.
Pretraining starts from unlabelled molecular data and requires semantic motif labels, atom and bond contextual vocabularies, and a split dataset structure. The documented preparation process produces feature arrays, vocabulary dictionaries and separate graph-data partitions. Finetuning uses a CSV containing a smiles column and can incorporate additional molecular descriptors stored in .npz files. A finetuned checkpoint then supports predictions written to an output file; predictions are averaged across the configured folds and ensemble members.
For representation-based workflows, GROVER exports fingerprints derived from pooled atom embeddings, pooled bond embeddings, or their concatenation. Additional molecular features can be appended when supplied. This makes the project relevant both to task-specific prediction and to evaluating learned molecular representations in separate downstream analyses.
The README identifies limitations relevant to reproduction: nondeterministic behavior associated with index_select_nd makes exact finetuning reproduction difficult, and this implementation adds a dense connection in the message-passing layer relative to the original implementation. It supplies evaluation procedures and finetuned checkpoints rather than establishing reproducibility on a new user's system. The project also explicitly states that it is not an officially supported Tencent product.
Key Features
- Self-supervised `GTransformer` pretraining using semantic motif labels and atom/bond contextual vocabularies, with documented single-GPU and Horovod-based multi-GPU workflows.
- Downloadable GROVER base and large pretrained checkpoints, plus finetuned checkpoints for eleven molecular benchmark datasets.
- Finetuning on labelled CSV datasets containing a `smiles` column, with optional `rdkit_2d_normalized` molecular features stored as `.npz` arrays.
- Prediction from finetuned models, averaging outputs across the configured number of folds and ensemble members.
- Molecular fingerprint export using atom embeddings, bond embeddings, or their concatenation, with optional additional molecular features.
- Checkpoint evaluation with documented AUC settings for classification, RMSE for regression by default, and MAE for QM7 and QM8 reproduction.
Use Cases
- Suggested evaluation: adapt a pretrained GROVER checkpoint to a labelled molecular classification or regression dataset and assess performance on a held-out split.
- Suggested evaluation: compare atom, bond and combined GROVER fingerprints as inputs to a separate downstream property-prediction analysis.
- Suggested evaluation: use the supplied finetuned benchmark checkpoints and documented metrics to investigate agreement with the README's reported experiments.
- Suggested evaluation: pretrain molecular representations on an unlabelled collection after preparing motif labels, contextual vocabularies and data partitions.
How to Use
- Read the official README and choose between pretraining, finetuning, prediction, fingerprint generation and evaluation. Inspect its requirements and environment instructions; it specifies Python 3.6.8 and references
requirements.txtand aDockerfile. - For a pretrained starting point, obtain the GROVER base checkpoint or select the large-model link in the README. Keep pretrained representation checkpoints distinct from task-specific finetuned models.
- Prepare inputs following the relevant README example. Finetuning requires a CSV with
smiles; pretraining additionally requires motif labels, atom/bond vocabularies and dataset splitting, including for single-GPU runs. - Follow the documented finetuning or fingerprint workflow. If training used additional molecular features, generate matching features for prediction inputs. Choose atom, bond or combined fingerprints according to the intended analysis.
- Evaluate saved checkpoints using the documented task metric and split settings. Treat this as a proposed local evaluation, account for the stated nondeterminism and dense-connection difference, and consult the code licence separately from checkpoint or dataset terms.