Define the planning problem
Start with the decision the evaluation must support: selecting a route-planning toolkit, reproducing a single-step prediction workflow, or preparing reaction data. These are different tasks and should not share a single undifferentiated leaderboard.
Specify target structures, permitted starting materials, chemistry-domain boundaries and the intended use of outputs. Record whether reaction categories are available, whether stereochemistry matters and what constitutes an unsupported case. A generated precursor set is not a complete route; a complete-looking route is not an experimentally validated procedure.
Before running anything, define independent acceptance criteria. Can the output be interpreted? Does a route terminate in the configured inventory? Is the result reproducible under the recorded settings? Which questions require a qualified chemist rather than an automated score?
Match each resource to its documented role
The linked resource references describe four distinct roles:
- AiZynthFinder is a retrosynthetic planning toolkit. Its default approach uses Monte Carlo tree search and a neural-network expansion policy based on reaction templates. Search works toward precursors in a configured stock. Alternative search algorithms and expansion policies are also documented.
- LocalRetro provides a dataset-based retrosynthesis research workflow using local reaction templates. Its stages include template extraction, data labeling, model training, test prediction and decoding into candidate reactants.
- RetroXpert separates product bond-disconnection prediction from reactant generation. The documented pipeline produces synthons from predicted disconnections and uses an OpenNMT-based second stage.
- RxnMapper assigns atom correspondences to supplied reaction SMILES. It does not predict products from reactants or search for synthesis routes.
Evaluate each role separately. Connecting these resources would be a proposed integration, not evidence that their representations, dependencies or outputs already interoperate.
Establish an input/output contract
Create a short specification for each evaluation track before installation or model execution:
| Track | Source-described inputs | Source-described outputs | Evaluation question |
|---|---|---|---|
| AiZynthFinder planning | Target SMILES, configuration, stock and expansion policy; optional filter policy | Candidate retrosynthetic routes | Do routes reach the configured stock under the chosen search settings? |
| LocalRetro reproduction | Reaction datasets and extracted local templates | Checkpoint, raw test predictions and decoded reactants | Are preprocessing and scoring conventions reproduced consistently? |
| RetroXpert reproduction | Reaction data processed into labels and graphs, followed by synthon preparation | Disconnection predictions and generated reactants | Which stage contributes to an observed failure? |
| RxnMapper data preparation | Valid reaction SMILES | Atom-mapped reactions and confidence values | Are mappings usable, and are invalid outputs detected? |
LocalRetro's linked references describe dataset training and testing, not an established custom-target interface. Do not silently turn that workflow into an interactive service. Similarly, mapping a complete reaction requires information that would not be available when only a target molecule is supplied.
Inspect data and preprocessing
Record reaction-data provenance, identifiers, preprocessing decisions and the terms applicable to each asset. Keep code, weights, stocks and datasets separate in the permissions record; a repository licence alone does not settle every asset's permitted use.
Check atom mapping, reaction classes, stereochemistry handling and split construction. Ask whether the evaluation resembles the intended chemistry domain and whether information available only in the reference reaction could influence a prediction.
RetroXpert deserves a specific preprocessing check: its reference describes information leakage through atom ordering and mapping numbers during synthon preparation, with revised product canonicalization and mapping reassignment recommended. Record the implementation used rather than assuming announced updates were completed.
For LocalRetro, state whether scoring is stereo-aware or stereo-unaware. For RxnMapper batch processing, treat >> or an empty dictionary as failure outputs, not successful mappings. Its numerical confidence is not established here as a calibrated probability of correctness.
Work through a hypothetical evaluation plan
Hypothetical example: a team wants to assess whether route suggestions would help experts triage a small internal target list. It selects 12 targets spanning three relevant chemistry groups and freezes a starting-material inventory before any runs. These numbers are planning choices, not reported results.
- Preserve the original targets and document any representation changes.
- Configure AiZynthFinder with a named stock and expansion policy, then record search settings and available outputs.
- Assess LocalRetro and RetroXpert separately on a bounded reference-reaction subset using their documented dataset workflows. Do not score them as complete route planners.
- If RxnMapper is proposed for reference-data preparation, retain both original and mapped reactions and flag mapping failures. Do not feed reference reactants into a target-only planning task.
- Have a qualified reviewer classify suggestions as requiring further assessment, unsupported or uninterpretable, with reasons.
The decision is whether to continue a larger evaluation—not whether any proposed route is ready for laboratory execution. Report missing outputs and failed runs alongside interpretable cases.
Separate execution from scientific quality
Maintain two records. The execution record covers environment details, dependency resolution, model and data access, configuration, input files and output artifacts. The scientific record covers precursor plausibility, route assumptions, scoring conventions and reviewer concerns.
Use AiZynthFinder's documentation, the LocalRetro README, the RetroXpert README and the RxnMapper README for workflow details. Documented environments are starting points, not proof of current compatibility. Successful execution establishes neither chemical feasibility nor comparative accuracy.
Interpret routes with human review
Keep stock membership, ranking, depth limits and other search assumptions visible. A high-ranked route may depend on unsuitable availability assumptions or omit conditions needed to assess a transformation. Agreement with a reference reactant set also does not establish yield, selectivity or safety.
The available evidence does not establish comparative performance, confidence calibration or laboratory validation. Require independent expert consideration of hazards, materials and procedures before any experimental decision. Revisit the evaluation when models, datasets, stocks or selection rules change.
Checklist and further discovery
- Define the task and freeze the evaluation inputs.
- Record assets, preprocessing and scoring conventions.
- Keep mapping, single-step prediction and route planning distinct.
- Capture invalid outputs, failures and unresolved assumptions.
- Separate execution success from chemistry assessment.
- Require a documented human decision before further use.
Browse resources and browse scientific tasks to identify related options and source links. This guide proposes an evaluation process; it does not certify performance, compatibility or experimental readiness.