Overview
ChemDataExtractor v2 is a toolkit for extracting chemical information from scientific literature. It supports a literature-processing workflow in which documents are read, their text is analysed with chemistry-aware natural language processing, and chemical information is extracted for subsequent use. It is an extraction toolkit rather than a literature database or a predictive chemistry model.
The documented input formats are HTML, XML and PDF. Its extraction components include chemical named entity recognition, rule-based grammars for properties and spectra, and a parser for tabulated data. These capabilities address both narrative text and tables, allowing users to evaluate different routes for recovering information from papers. Document-level processing is also described as resolving interdependencies among extracted data, although the project README does not explain the resolution rules or provide worked examples.
The intended outputs are extracted chemical information, including entities, property and spectra data, and table content. The source excerpts do not specify output schemas, export formats, supported property categories or extraction accuracy. Consequently, suitability for a particular literature collection should be evaluated using representative documents and manually checked results rather than assumed from the feature list.
The README provides a pip installation route and states support for Python 3.9 through Python 3.11. It also links to documentation for development and contribution guidance. The package metadata declares a command-line entry point named cde, but the excerpts do not document its invocation options. Python-version classifiers in setup.py differ from the README’s support statement, so they should not be treated as evidence of tested compatibility.
Key Features
- Document readers for HTML, XML and PDF scientific literature.
- Chemistry-aware natural language processing pipeline for document text.
- Chemical named entity recognition.
- Rule-based parsing grammars for extracting properties and spectra.
- Table parsing for recovering tabulated data.
- Document-level processing to resolve interdependencies among extracted data.
Use Cases
- Suggested evaluation: extract chemical entities from a representative set of papers and compare the results with manually annotated text.
- Suggested evaluation: recover property or spectra information from literature passages using the documented rule-based parsing capabilities.
- Suggested evaluation: extract data from scientific tables and check whether values and their contextual associations are preserved.
- Suggested evaluation: compare extraction from HTML, XML and PDF versions of representative documents before selecting an ingestion workflow.
How to Use
- Read the official README to confirm the documented input formats, extraction components and Python support range before choosing an environment.
- Create and activate an isolated environment using the README’s installation guidance. It gives a conda environment with Python 3.11 as an example; support is stated for Python 3.9–3.11, not independently verified here.
- Install the package with the supplied instruction,
pip install chemdataextractor2. Keep the package name distinct from the project’s display name, ChemDataExtractor. - Consult the linked documentation for workflow details. Determine how to connect the relevant document reader, text or table extraction components, and output handling; those details are not included in the source excerpts.
- As an evaluation step, process a small representative document set and manually inspect chemical entities, properties, spectra and table data against the originals. Record missing or incorrect extractions before considering a larger literature-processing workflow.