Skip to content

ChemDataExtractor

Matt Swain, Callum Court, Juraj Mavracic, Taketomo Isazawa, and contributors

ChemDataExtractor v2 is a Python toolkit for extracting chemical entities, properties, spectra and tabulated data from scientific literature, with readers for HTML, XML and PDF.

Catalog updated ·

Overview

ChemDataExtractor v2 is a toolkit for extracting chemical information from scientific literature. It supports a literature-processing workflow in which documents are read, their text is analysed with chemistry-aware natural language processing, and chemical information is extracted for subsequent use. It is an extraction toolkit rather than a literature database or a predictive chemistry model.

The documented input formats are HTML, XML and PDF. Its extraction components include chemical named entity recognition, rule-based grammars for properties and spectra, and a parser for tabulated data. These capabilities address both narrative text and tables, allowing users to evaluate different routes for recovering information from papers. Document-level processing is also described as resolving interdependencies among extracted data, although the project README does not explain the resolution rules or provide worked examples.

The intended outputs are extracted chemical information, including entities, property and spectra data, and table content. The source excerpts do not specify output schemas, export formats, supported property categories or extraction accuracy. Consequently, suitability for a particular literature collection should be evaluated using representative documents and manually checked results rather than assumed from the feature list.

The README provides a pip installation route and states support for Python 3.9 through Python 3.11. It also links to documentation for development and contribution guidance. The package metadata declares a command-line entry point named cde, but the excerpts do not document its invocation options. Python-version classifiers in setup.py differ from the README’s support statement, so they should not be treated as evidence of tested compatibility.

Key Features

  • Document readers for HTML, XML and PDF scientific literature.
  • Chemistry-aware natural language processing pipeline for document text.
  • Chemical named entity recognition.
  • Rule-based parsing grammars for extracting properties and spectra.
  • Table parsing for recovering tabulated data.
  • Document-level processing to resolve interdependencies among extracted data.

Use Cases

  • Suggested evaluation: extract chemical entities from a representative set of papers and compare the results with manually annotated text.
  • Suggested evaluation: recover property or spectra information from literature passages using the documented rule-based parsing capabilities.
  • Suggested evaluation: extract data from scientific tables and check whether values and their contextual associations are preserved.
  • Suggested evaluation: compare extraction from HTML, XML and PDF versions of representative documents before selecting an ingestion workflow.

How to Use

  1. Read the official README to confirm the documented input formats, extraction components and Python support range before choosing an environment.
  2. Create and activate an isolated environment using the README’s installation guidance. It gives a conda environment with Python 3.11 as an example; support is stated for Python 3.9–3.11, not independently verified here.
  3. Install the package with the supplied instruction, pip install chemdataextractor2. Keep the package name distinct from the project’s display name, ChemDataExtractor.
  4. Consult the linked documentation for workflow details. Determine how to connect the relevant document reader, text or table extraction components, and output handling; those details are not included in the source excerpts.
  5. As an evaluation step, process a small representative document set and manually inspect chemical entities, properties, spectra and table data against the originals. Record missing or incorrect extractions before considering a larger literature-processing workflow.

Related resources

ChemCrow is a Python chemistry agent package built with Langchain that connects language-model reasoning to RDKit, paper-qa, chemical databases, and reaction-planning tools.

Open sourcePython

Cheminformatics · Literature & Research

Coscientist provides a simple implementation and supporting research data for large-language-model chemical research, covering synthesis planning, library selections, repeated runs and optimization.

Literature & Research · Lab Automation

PaperQA

Agent

PaperQA provides agentic retrieval-augmented question answering over scientific documents, combining local search, ranked evidence summaries and answers with in-text citations.

Open sourcePython

Literature & Research

PubMed MCP (cyanheads) is a third-party MCP server connecting AI clients to biomedical literature search, article metadata, available full text, MeSH vocabulary and citation tools.

Open sourceTypeScript

Literature & Research · Drug Discovery

SciAgents is a research multi-agent framework that uses scientific knowledge graphs, LLMs and retrieval tools to develop and critique hypotheses for bio-inspired materials research.

Open sourcePython

Literature & Research · Materials Discovery

STELLA

Agent

STELLA is a research multi-agent framework for biomedical literature and bioinformatics workflows, combining reusable workflow templates, tool orchestration, critique, and optional dynamic tool creation.

Open sourcePython

Literature & Research · Drug Discovery

Related guides

Agents

AI Agents for Chemistry: From Chatbots to Autonomous Research

A practical guide to eight chemistry and biomedical research agents: how tools, memory, planning and multi-agent roles work, which projects fit different tasks, and how to evaluate bounded autonomy with scientific oversight.

Workflows

AI Literature Research for Chemistry

Build a traceable chemistry literature workflow with PubMed MCP: refine searches, inspect metadata and abstracts, synthesize multiple papers, follow citations, and separate retrieved evidence from agent-generated hypotheses.