Start with a provenance record

Before combining chemistry datasets, establish what each input represents and where it came from. A download location alone is not enough: a mirror, processed archive and original release can contain related but different records.

Create one manifest entry per input artifact. Record the publisher, dataset name, release identifier, source URL, access date, filename and checksum. Include the publication date when available; mark missing information as unknown rather than inferring it. Identify whether the artifact is an original release or a derivative, and retain its citations alongside the data.

Input: downloaded files, source documentation and available release metadata. Output: a manifest linking each artifact to its origin and any known upstream releases.

Decision questions: Can someone retrieve the same artifact? Can each processed row be traced to an original record? If a release identifier is unavailable, what evidence distinguishes this download from a later one?

Separate data rights from code licensing

Inspect the dataset licence and access terms independently of the repository licence. Record the relevant text or its location, what material it applies to, and unresolved questions about redistribution, commercial use and derivatives. Do not assume that publicly downloadable data carry permission for every intended use.

The linked resource references illustrate why this distinction matters:

  • Therapeutics Data Commons describes an MIT code licence, while dataset usage rights require individual checks.
  • Open Reaction Database distinguishes dataset licensing from repository code licensing.
  • SPICE describes dataset usage as CC0, separately from its repository’s MIT licence. Record that source statement for the selected release rather than extending it to unrelated inputs.

For every planned operation, ask: Are we permitted to download, transform, train on and redistribute this material? If the evidence does not answer a question, document the uncertainty and seek clarification. A provenance checklist is not a legal determination.

Validate identifiers, units and scientific meaning

Preserve source identifiers and original values before creating normalized fields. Keep a mapping between raw and processed records, including unsuccessful conversions and ambiguous matches.

Inspect units, missing-value conventions, experimental conditions and calculation methods. A shared property name does not establish that labels are interchangeable. Likewise, matching molecular connectivity does not establish that two conformations or measurements are duplicates.

Resource-specific details can affect the mapping. Open Reaction Database stores serialized reaction records in Parquet files, and its reference identifies the reaction_id column as authoritative. It also describes retired_datasets.csv for mapping consolidated dataset IDs. Preserve these relationships when processing older references.

For computed labels, retain the program, theoretical method and available settings. SPICE explicitly cautions that matching the level of theory alone is insufficient when generating compatible reference calculations: the program and settings must also match.

Output: a schema dictionary, identifier crosswalk and validation report listing unresolved records rather than silently overwriting them.

Check representation, coverage and duplication

Choose the unit of analysis before counting records: molecule, conformer, reaction, measurement or molecular cluster. Then inspect coverage relevant to the intended task, such as elements, charge states, molecular classes, target ranges and measurement protocols. Consider demographic coverage where the data and task make it relevant.

GEOM provides molecular conformations with energy and statistical-weight annotations. Its reference notes that MessagePack archives are not updated with additions made to the Python-specific data. Record the representation used; the resource name alone does not establish identical coverage.

Check exact duplicates, normalized-structure matches and correlated observations. Multiple conformers of one molecule or measurements drawn from a shared source can leak information across partitions even when row identifiers differ. Choose grouping rules appropriate to the scientific question, and retain ambiguous cases for inspection.

For Matbench, the reference establishes a suite of 13 materials-science tasks, but not their individual input schemas or split details. Consult the linked documentation before treating a task as compatible with another dataset.

Version every transformation

Keep parsing code, configuration, checksums and upstream release identifiers with each processed release. Record filtering rules, unit conversions, identifier normalization, label transformations and partition assignments. Preserve raw inputs separately so that a revised rule does not erase the original evidence.

For each stage, capture input and output record counts and reasons for exclusions. Counts help locate processing errors; they do not demonstrate scientific correctness. Make manual decisions reproducible through a small decision log tied to source identifiers.

Therapeutics Data Commons describes loaders, processing functions and configurable splits, including a scaffold-split example. These capabilities can support preparation, but users still need to record the selected dataset, parameters and returned partitions. A documented feature is not evidence that a particular workflow has been tested successfully.

Worked planning example: comparing conformer sources

Hypothetical scenario: a team wants to explore whether GEOM and SPICE could support a molecular-energy learning study. This is a proposed evaluation plan, not a tested integration.

  1. Select one concrete artifact from each resource and record its release, representation, checksum and usage-condition evidence.
  2. Inspect a small sample to identify molecule keys, conformer coordinates, energy fields, units and calculation provenance.
  3. Build a crosswalk that keeps molecular identity separate from conformer identity. Flag possible overlapping molecules without assuming their geometries match.
  4. Keep the two sources’ labels separate until their definitions, reference conventions and computational settings have been assessed.
  5. Propose molecule-grouped partitions and quantify overlap before training. Reconsider the split if related records cross partition boundaries.

The planning output is a compatibility table and a go/no-go decision, not a merged training set. If label comparability remains unresolved, evaluate sources separately or narrow the question. Do not pool energies simply because both resources provide them.

Communicate limitations and unresolved decisions

Publish a data statement explaining origins, usage-condition evidence, transformations, exclusions, split logic and known coverage gaps. Include citations and a route for corrections. Distinguish source-described capabilities from checks you propose to perform and checks actually completed by your team.

Provenance tracking does not certify label accuracy, model performance or fitness for clinical or other consequential use. Sparse observations, differing protocols and uncertain rights remain limitations even when the processing pipeline is reproducible.

Actionable checklist and next steps

  • Record each artifact’s origin, release and checksum.
  • Separate dataset terms from code licences.
  • Preserve raw identifiers, values and calculation context.
  • Review ambiguous matches and correlated records before splitting.
  • Version transformations and document unresolved decisions.

Browse resources and browse scientific tasks to identify candidate inputs and evaluation questions. Follow source links for release-specific details; this guide does not certify scientific performance or compatibility.