Skip to content
Community

A QM9 Dipole GNN Teaching Baseline with PyTorch Geometric

Prepare QM9 molecular graphs, train a small CPU GCN baseline with training-only normalization, and report validation-selected test MAE in Debye.

Level: Intermediate Cost: Free Privacy: Local ~60 min
Start Setup

You'll be able to

  • Produce traceable computational candidate artifacts and inspect the documented acceptance criteria.

What you'll build

Documentation-based review draft. This workflow has not been executed in this batch; no installation, inference, optimization or experimental result has been verified.

Audience and scope

For learners of molecular graph regression. Target 0 is QM9 dipole moment, unit Debye (D). The reviewed baseline uses two GCN layers, atom features and graph mean pooling; it intentionally ignores bond-type edge attributes and 3D coordinates. It is not a state-of-the-art model or a geometry-aware dipole estimator.

Data and permissions

The original Figshare data article 1057646 explicitly declares CC0, independently of the paper license and PyG MIT code license. Preserve the original DOI, article version and metadata. The pinned PyG QM9 source uses the DeepChem gdb9.zip mirror and Figshare uncharacterized file 3195404 with RDKit, excludes the listed uncharacterized molecules, and exposes 19 targets. Require RDKit and a fresh root to avoid silently reusing the alternative no-RDKit processed archive. Record the actual sample count and SHA-256 of downloaded and processed files; this batch has not downloaded the dataset or established byte equivalence to the original XYZ archive. Do not add extra filters silently.

Training contract and outputs

Seed 37 creates disjoint 8000/1000/1000 teaching splits from a permutation; unused samples are deliberate, not a full benchmark. Fit the target mean/std only to training labels. Ten CPU epochs use Adam and normalized MSE; choose the checkpoint using validation MAE and evaluate test MAE once in D. The upstream NNConv example normalizes all labels before its split, so it is not copied unchanged. Expected files: split indices, checkpoint, dataset hashes, per-epoch training/validation logs and results.json. No score is claimed before execution.

Limits and costs

Random in-distribution splitting does not demonstrate scaffold or external generalization. No hyperparameter tuning on test data. Compare against a training-only constant baseline and then separate scaffold/OOD evaluation in later work. Python >=3.10 is an upstream requirement; this review selects Python 3.11/Linux/CPU without a tested dependency lock. Installation, downloads, memory and training consume resources. Dataset downloads access public external hosts; molecular data and training remain local. Retain PyG MIT notices and original data attribution.

Official references

Molecular graph dataset and GCN training

PyTorch Geometric

Stack Components

PyTorch Geometric

Molecular graph dataset and GCN training · source 79d33965a40b7fa83616a9f598a0f8619f25d939; Python 3.11; RDKit raw route

Documentation-based review draft. This workflow has not been executed in this batch; no installation, inference, optimization or experimental result has been verified.

Code is open source; compute/storage and any external-service conditions remain the user's responsibility.

View Resource

Compatibility

ClientOSArchitectureVersion requirements
Python LinuxAny>= 3.11

Setup & Test

1. Specify the dipole task and dataset rights

Linux

Use target 0 in D. Preserve the original CC0 data metadata and the pinned QM9 source; inspect its raw download and uncharacterized-molecule filter policy.

Official source

Expected result

A documented target, data origin and processing policy.

2. Prepare a versioned CPU environment

Linux

Use Python 3.11, a compatible CPU PyTorch build, NumPy and RDKit. Follow official PyG installation guidance, install from the inspected commit and export resolved dependencies; no tested lock is claimed.

git clone https://github.com/pyg-team/pytorch_geometric.git
git -C pytorch_geometric checkout 79d33965a40b7fa83616a9f598a0f8619f25d939
python -m pip install ./pytorch_geometric
Official source

Expected result

Recorded environment with importable QM9/GCNConv/DataLoader.

3. Review training-only preprocessing and splits

Linux

Save the accompanying authored code as qm9_gcn_baseline.py. It requires RDKit and a fresh data root, records actual dataset hashes, fixes seed/splits, and derives mean/std only from training target 0. It is a composition of documented APIs, not an upstream reproduction.

"""Review-only authored baseline using documented PyG APIs; never run in this batch."""
import argparse
import copy
import hashlib
import json
from pathlib import Path
import random
import numpy as np
import rdkit  # Require the inspected raw-processing route, not the no-RDKit archive.
import torch
from torch import nn
import torch_geometric
from torch_geometric.datasets import QM9
from torch_geometric.loader import DataLoader
from torch_geometric.nn import GCNConv, global_mean_pool

p = argparse.ArgumentParser()
p.add_argument('data_root')
p.add_argument('output')
a = p.parse_args()
out = Path(a.output)
out.mkdir(parents=True, exist_ok=False)
root = Path(a.data_root)
if root.exists() and any(root.iterdir()):
    raise ValueError('Use a fresh data root to prevent an undocumented processed cache')
seed = 37
random.seed(seed)
np.random.seed(seed)
torch.manual_seed(seed)
torch.use_deterministic_algorithms(True)
torch.set_num_threads(1)
dataset = QM9(str(root))
# The inspected raw route removes uncharacterized molecules; record actual count.
if len(dataset) < 10000:
    raise ValueError('Unexpectedly small QM9 dataset')
ids = torch.randperm(len(dataset), generator=torch.Generator().manual_seed(seed))
# A deliberately small teaching run, not a literature benchmark split.
train_ids, val_ids, test_ids = ids[:8000], ids[8000:9000], ids[9000:10000]
assert not (set(train_ids.tolist()) & set(val_ids.tolist()))
assert not (set(train_ids.tolist()) & set(test_ids.tolist()))
assert not (set(val_ids.tolist()) & set(test_ids.tolist()))
targets = torch.cat([dataset[int(i)].y[:, 0] for i in train_ids])
mu, sigma = targets.mean(), targets.std(unbiased=False)
if not torch.isfinite(targets).all() or not torch.isfinite(sigma) or sigma <= 0:
    raise ValueError('Invalid training targets')
loaders = [DataLoader(dataset[ix], batch_size=64, shuffle=(j == 0),
    num_workers=0, generator=torch.Generator().manual_seed(seed))
    for j, ix in enumerate([train_ids, val_ids, test_ids])]

class Baseline(nn.Module):
    def __init__(self):
        super().__init__()
        self.c1 = GCNConv(dataset.num_node_features, 64)
        self.c2 = GCNConv(64, 64)
        self.head = nn.Linear(64, 1)

    def forward(self, batch):
        x = self.c1(batch.x.float(), batch.edge_index).relu()
        x = self.c2(x, batch.edge_index).relu()
        return self.head(global_mean_pool(x, batch.batch)).flatten()

model = Baseline()  # CPU only; no geometry or edge_attr features in this baseline.
optimizer = torch.optim.Adam(model.parameters(), lr=0.001)

@torch.no_grad()
def mae(loader):
    model.eval()
    total, count = 0.0, 0
    for batch in loader:
        predicted = model(batch) * sigma + mu
        errors = (predicted - batch.y[:, 0]).abs()
        if not torch.isfinite(errors).all():
            raise ValueError('Nonfinite evaluation')
        total += errors.sum().item()
        count += batch.num_graphs
    return total / count  # Debye; no normalization statistics fitted to this split.

history, best, weights = [], float('inf'), None
for epoch in range(1, 11):
    model.train()
    loss_sum, n = 0.0, 0
    for batch in loaders[0]:
        optimizer.zero_grad()
        loss = nn.functional.mse_loss(model(batch), (batch.y[:, 0] - mu) / sigma)
        if not torch.isfinite(loss):
            raise ValueError('Nonfinite loss')
        loss.backward()
        optimizer.step()
        loss_sum += loss.item() * batch.num_graphs
        n += batch.num_graphs
    val = mae(loaders[1])
    history.append({'epoch': epoch, 'train_normalized_mse': loss_sum / n, 'val_mae_D': val})
    if val < best:
        best, weights = val, copy.deepcopy(model.state_dict())
model.load_state_dict(weights)
test_mae = mae(loaders[2])  # Test once after validation-only selection.
torch.save(weights, out / 'best_state_dict.pt')
np.savez(out / 'splits.npz', train=train_ids.numpy(), validation=val_ids.numpy(), test=test_ids.numpy())
files = {str(f.relative_to(root)): hashlib.sha256(f.read_bytes()).hexdigest()
    for f in root.rglob('*') if f.is_file()}
report = {'target': 0, 'property': 'dipole moment', 'unit': 'D', 'seed': seed,
    'dataset_count': len(dataset), 'split_sizes': [8000, 1000, 1000],
    'train_mean_D': mu.item(), 'train_std_D': sigma.item(), 'epochs': 10,
    'validation_mae_D': best, 'test_mae_D': test_mae, 'history': history,
    'versions': {'torch': torch.__version__, 'pyg': torch_geometric.__version__,
        'rdkit': rdkit.__version__, 'numpy': np.__version__},
    'device': 'cpu', 'dataset_file_sha256': files,
    'scope': 'Teaching baseline; small random split, no bond types or coordinates; no transfer guarantee.'}
(out / 'results.json').write_text(json.dumps(report, indent=2) + '\n')
Official source

Expected result

Disjoint teaching splits and train-only normalization are explicit.

4. Train and evaluate without test selection

Linux

Run in the recorded environment with new data and output directories. Inspect raw graph shapes, finite features/targets and the dataset manifest; select by validation only and report the single final test MAE in D.

python qm9_gcn_baseline.py qm9_fresh qm9_review
Official source

Expected result

Expected: splits.npz, best_state_dict.pt and results.json with per-epoch logs, unit-labelled metrics and dataset hashes. No observed metrics in this batch.

Troubleshooting

  • RDKit missing or old cache: create the required environment and fresh root.
  • Nonfinite labels/loss: stop and inspect data/units, never replace silently.
  • Unexpected count: inspect exclusion list and file hashes.
  • Poor MAE: compare to train-mean baseline and inspect splits; test data must not tune the model.
Still not working

Alternatives

Official NNConv is a separate bond/geometry-aware direction; correct its whole-dataset normalization before adaptation. This GCN baseline makes no quality equivalence claim.