Overview
Open Reaction Database (ORD), distributed through ord-data, is a reaction-data repository rather than an installable package or a predictive model. It supplies structured records for downstream chemistry data workflows, while helper scripts support downloading and processing datasets. GitHub remains the authoritative repository; Hugging Face serves as a read mirror for dataset objects, including downloads routed automatically through Git LFS.
Each Parquet file represents one dataset, with serialized reaction Protobuf messages stored row by row and the dataset name, description and ID retained in file metadata. The documented ord_schema.datasets.load_dataset interface returns a streaming DatasetView by default, allowing reactions to be read on demand. Users can iterate, index or slice records, retrieve a reaction by ID, enumerate IDs, or process individual row groups. Materializing a Dataset proto enables whole-dataset mutation and serialization; examples also demonstrate human-readable .pbtxt output and conversion of individual reactions to JSON.
Important representation limits affect integration. Dataset.reaction_ids is not persisted, the reaction_id column is authoritative, and empty datasets cannot be written. Indexed reactions are newly deserialized copies, so changing them does not alter the stored file. The similarly named ord_schema.parquet.load_dataset always materializes data rather than streaming it. For older references, retired_datasets.csv maps consolidated dataset IDs, while USPTO reaction records retain patent provenance. Dataset licensing and repository code licensing are separate.
The README assigns Apache-2.0 to repository helper code and CC-BY-SA-4.0 to datasets and their descriptive metadata. Check these components separately, including attribution and redistribution conditions for the data you use.
Key Features
- Stores one dataset per Parquet file, with serialized reaction Protobuf records and dataset-level name, description and ID metadata.
- Supports full Git LFS downloads through an automatically configured Hugging Face mirror, plus data-only and subset downloads using the supplied helper.
- Provides documented streaming access through ord_schema, including iteration, indexing, slicing, reaction-ID lookup and ID-only enumeration.
- Exposes row-group iteration yielding `(reaction_id, Reaction)` pairs for partitioning work across large files.
- Documents conversion to human-readable `.pbtxt` and conversion of individual reaction messages to JSON.
- Maps retired dataset IDs to consolidated replacements and retains per-reaction patent provenance for USPTO data.
Use Cases
- Intended evaluation: inspect representative reaction records before deciding whether a selected dataset is suitable for a reaction-prediction training or benchmarking workflow.
- Intended evaluation: build a row-group-based extraction pipeline and check whether streaming access meets local memory and processing requirements.
- Intended evaluation: convert selected records to JSON or `.pbtxt` for schema inspection and integration with downstream chemistry data tools.
- Intended evaluation: reconcile older dataset references with consolidated IDs and inspect USPTO patent provenance for traceability.
How to Use
- Read the repository README and identify the datasets needed. Treat
ord-dataas a data repository, not a package to install. - Choose full-repository access with Git LFS or browse the Hugging Face mirror. For selective downloads, follow the README’s helper-script instructions rather than fetching the entire corpus.
- Follow the documented
ord_schemaloading example on one selected Parquet file. Begin with the default streaming view and inspect its metadata, record count and representative reactions. - Select iteration, ID lookup or row-group processing for your workflow. Materialize a Dataset proto only when whole-dataset mutation or serialization is required; use the documented examples for
.pbtxtor JSON output. - Before redistribution or submission, review the dataset license, terms of use and, where relevant, the Submission Workflow.