Overview
MolScribe recognizes chemical structures from molecular images using an image-to-graph model. It serves the structure-recognition stage of a chemistry document-processing workflow: a molecular diagram is supplied as input, and the prediction interface returns structured chemical representations. The repository provides a Python package, downloadable model checkpoints, prediction examples, and training and evaluation resources. A linked HuggingFace demo offers another way to explore the model.
The documented prediction API returns a dictionary containing SMILES and molfile representations. Optional outputs include an overall confidence value, atom symbols and coordinates, bond types and endpoint indices, and atom- and bond-level confidence values. The quick-start example loads a checkpoint and predicts from an image file on a CPU. The architecture uses a Swin-B image encoder and a six-layer Transformer decoder, with a documented input size of 384×384.
For research workflows, the README links training datasets derived from USPTO and PubChem, alongside synthetic, realistic and perturbed image benchmarks. It also describes training scripts and a standalone evaluator that compares prediction CSV files with reference structures. Evaluation distinguishes canonical SMILES exact matching, graph exact matching that ignores tetrahedral chirality, and exact matching on chiral molecules.
These resources support assessing recognition on a chosen image collection, but the source excerpts do not establish accuracy for a particular deployment or document type. MolScribe is presented as molecular-image recognition; reaction-diagram parsing and broader literature extraction are linked as separate projects. The repository's MIT code licence should not be treated as evidence of checkpoint or dataset terms.
Key Features
- Image-to-graph prediction of chemical structures from molecular image files, with SMILES and molfile outputs.
- Optional atom and bond records, including atom coordinates, bond types and endpoint atom indices.
- Optional confidence reporting for the overall prediction and individual atoms and bonds.
- Downloadable checkpoints and prediction examples through a Python interface, script and Jupyter notebook.
- Training and evaluation resources covering USPTO and PubChem training data plus synthetic, realistic and perturbed benchmarks.
- Standalone CSV-based evaluation reporting canonical SMILES, graph and chiral exact-match scores.
Use Cases
- Intended evaluation: convert molecular diagram images from a literature-processing pipeline into candidate SMILES and molfiles for subsequent checking.
- Intended evaluation: use atom coordinates and bond connectivity to inspect recognition errors against the source diagram.
- Intended evaluation: compare predictions with reference structures across synthetic, realistic or perturbed image collections using the documented evaluator.
- Intended evaluation: reproduce or extend the documented training workflow using the linked datasets and checkpoints.
How to Use
- Read the official README to choose between a quick prediction workflow and the separate training or evaluation setup. The linked demo is an alternative starting point.
- Follow a documented package-installation route. Check the declared Python requirement and dependencies in setup.py against your intended environment; these declarations are not proof of tested compatibility.
- Obtain a checkpoint from the HuggingFace Hub. Use the checkpoint associated with your chosen README example rather than assuming the quick-start and experiment examples are identical.
- Follow the README's image-file prediction example. Request atom/bond records and confidence information if needed, then inspect the returned SMILES, molfile and graph details against the input image.
- For intended evaluation, prepare prediction and reference CSV files with matching
image_idvalues. Follow the README's standalone evaluation instructions, select the prediction column as documented, and interpret the three exact-match metrics separately.