Skip to content

Open Reaction Database

Open Reaction Database Project Authors

Open Reaction Database provides structured reaction datasets as Parquet files containing Protobuf records, with mirrored downloads and ord_schema workflows for streaming access and text or JSON conversion.

Catalog updated ·

Overview

Open Reaction Database (ORD), distributed through ord-data, is a reaction-data repository rather than an installable package or a predictive model. It supplies structured records for downstream chemistry data workflows, while helper scripts support downloading and processing datasets. GitHub remains the authoritative repository; Hugging Face serves as a read mirror for dataset objects, including downloads routed automatically through Git LFS.

Each Parquet file represents one dataset, with serialized reaction Protobuf messages stored row by row and the dataset name, description and ID retained in file metadata. The documented ord_schema.datasets.load_dataset interface returns a streaming DatasetView by default, allowing reactions to be read on demand. Users can iterate, index or slice records, retrieve a reaction by ID, enumerate IDs, or process individual row groups. Materializing a Dataset proto enables whole-dataset mutation and serialization; examples also demonstrate human-readable .pbtxt output and conversion of individual reactions to JSON.

Important representation limits affect integration. Dataset.reaction_ids is not persisted, the reaction_id column is authoritative, and empty datasets cannot be written. Indexed reactions are newly deserialized copies, so changing them does not alter the stored file. The similarly named ord_schema.parquet.load_dataset always materializes data rather than streaming it. For older references, retired_datasets.csv maps consolidated dataset IDs, while USPTO reaction records retain patent provenance. Dataset licensing and repository code licensing are separate.

The README assigns Apache-2.0 to repository helper code and CC-BY-SA-4.0 to datasets and their descriptive metadata. Check these components separately, including attribution and redistribution conditions for the data you use.

Key Features

  • Stores one dataset per Parquet file, with serialized reaction Protobuf records and dataset-level name, description and ID metadata.
  • Supports full Git LFS downloads through an automatically configured Hugging Face mirror, plus data-only and subset downloads using the supplied helper.
  • Provides documented streaming access through ord_schema, including iteration, indexing, slicing, reaction-ID lookup and ID-only enumeration.
  • Exposes row-group iteration yielding `(reaction_id, Reaction)` pairs for partitioning work across large files.
  • Documents conversion to human-readable `.pbtxt` and conversion of individual reaction messages to JSON.
  • Maps retired dataset IDs to consolidated replacements and retains per-reaction patent provenance for USPTO data.

Use Cases

  • Intended evaluation: inspect representative reaction records before deciding whether a selected dataset is suitable for a reaction-prediction training or benchmarking workflow.
  • Intended evaluation: build a row-group-based extraction pipeline and check whether streaming access meets local memory and processing requirements.
  • Intended evaluation: convert selected records to JSON or `.pbtxt` for schema inspection and integration with downstream chemistry data tools.
  • Intended evaluation: reconcile older dataset references with consolidated IDs and inspect USPTO patent provenance for traceability.

How to Use

  1. Read the repository README and identify the datasets needed. Treat ord-data as a data repository, not a package to install.
  2. Choose full-repository access with Git LFS or browse the Hugging Face mirror. For selective downloads, follow the README’s helper-script instructions rather than fetching the entire corpus.
  3. Follow the documented ord_schema loading example on one selected Parquet file. Begin with the default streaming view and inspect its metadata, record count and representative reactions.
  4. Select iteration, ID lookup or row-group processing for your workflow. Materialize a Dataset proto only when whole-dataset mutation or serialization is required; use the documented examples for .pbtxt or JSON output.
  5. Before redistribution or submission, review the dataset license, terms of use and, where relevant, the Submission Workflow.

Related resources

Benchling MCP (longevity-genie) is a Python MCP server that connects AI clients to Benchling notebook entries, biological sequences, projects, and entity search using API credentials.

Open sourcePython

Lab Automation · Scientific Data

ChEMBL MCP (cyanheads) connects MCP clients to ChEMBL compound, target, bioactivity and drug records, with structure searches and optional DuckDB analysis of larger activity sets.

Open sourceTypeScript

Drug Discovery · Scientific Data

Official Python client maintained by the ChEMBL group for querying ChEMBL data and cheminformatics services, with QuerySet-style filters, lazy retrieval and local result caching.

Open sourcePython

Drug Discovery · Scientific Data

Academic materials data platform for exploring computed structures and properties and preparing materials-discovery research.

Materials Discovery · Scientific Data

Materials Project API is the Python client project published as mp-api, with official data-access documentation and an optional MCP server entry point declared in its package configuration.

Open sourcePython

Materials Discovery · Scientific Data

A third-party MCP server that connects assistant clients to Materials Project through mp_api, with tools for material searches, crystal structures, electronic properties, and other materials data.

Open sourcePython

Materials Discovery · Scientific Data

Related guides

Licensing

Reading Code, Weight and Data Licences Separately

Build a component-by-component licensing record for chemistry AI workflows, separating software, model weights, datasets and hosted services while keeping missing evidence and unresolved conditions visible.