Overview
GuacaMol provides a benchmarking framework for models that generate molecules, rather than a molecular generation model itself. It supports two evaluation roles: assessing whether generated molecules resemble a training distribution and assessing generation against scoring objectives. The repository also links standardized ChEMBL-derived datasets, making the resource relevant as both evaluation software and a source of benchmark data.
To connect a model to distribution-learning benchmarks, users implement DistributionMatchingGenerator and pass an instance, together with the training-set file location, to assess_distribution_learning. For goal-directed evaluation, users implement GoalDirectedGenerator, whose role is to produce a requested number of molecules with high scores under a supplied scoring function, and invoke assess_goal_directed_generation. These workflows yield benchmark assessments; the source excerpts do not specify the result-file format. A separate baseline repository provides example implementations and an example Docker environment.
The data workflow offers pre-built training, validation, test and combined datasets with published MD5 hashes. Alternatively, the repository includes a process for downloading and preparing ChEMBL data, excluding molecules very similar to a designated holdout set. The README warns that package-version differences can change generated datasets and cause preparation to fail, and describes a Docker-based route for reproducing the standardized data. RDKit and FCD are dependencies, with FCD used to calculate Fréchet ChemNet Distance. The supplied documentation establishes benchmark workflows, not prospective validation of generated molecules or evidence of any particular model's performance.
Key Features
- Distribution-learning evaluation through `DistributionMatchingGenerator` and `assess_distribution_learning`, using a model adapter and training-set file.
- Goal-directed evaluation through `GoalDirectedGenerator` and `assess_goal_directed_generation`, supporting molecule generation against supplied scoring functions.
- Pre-built ChEMBL-derived training, validation, test and combined datasets with published MD5 checksums.
- Dataset preparation that downloads and processes ChEMBL while excluding molecules very similar to the designated holdout set.
- Fréchet ChemNet Distance calculation through the FCD dependency.
- Links to baseline model implementations and Docker environments for benchmark development and standardized dataset preparation.
Use Cases
- Suggested evaluation: compare molecular generators on distribution-learning benchmarks using the same standardized training dataset.
- Suggested evaluation: assess a generator's ability to produce a requested number of molecules under goal-directed scoring objectives.
- Suggested evaluation: prepare a standardized ChEMBL-derived dataset and check its published hash before comparing benchmark results across environments.
How to Use
- Read the official README and choose distribution-learning or goal-directed evaluation according to the question you want to assess. Consult the linked benchmark paper for benchmark rationale and baseline-score context.
- Install the package using the README's installation guidance. Check dependencies against setup.py; dependency declarations differ from the README, so do not assume the two describe identical environments.
- Obtain the relevant splits from the GuacaMol data project and check the published MD5 hashes. If regenerating data, follow the README's preparation and Docker guidance.
- Implement
DistributionMatchingGeneratororGoalDirectedGeneratorfor your model. Use the baseline repository as an implementation reference rather than treating it as part of your model. - Invoke the corresponding assessment function, supplying the training-set location for distribution learning. Record the dataset identity and environment alongside results; the README also describes unit tests for checking installation.