2. Install isolated Python dependencies
macOS
Run in a new working directory. This shell variant applies to macOS/Linux; use the separate Windows step on Windows.
python3.11 -m venv .venv
.venv/bin/python -m pip install rdkit==2025.3.1 pubchempy==1.0.5
.venv/bin/python -m pip freeze
Official sourceExpected result
Packages install; record the resolved versions.
3. Install isolated Python dependencies
Linux
Run in a new working directory. This shell variant applies to macOS/Linux; use the separate Windows step on Windows.
python3.11 -m venv .venv
.venv/bin/python -m pip install rdkit==2025.3.1 pubchempy==1.0.5
.venv/bin/python -m pip freeze
Official sourceExpected result
Packages install; record the resolved versions.
4. Install Python dependencies on Windows
Windows
Use PowerShell and the virtual-environment interpreter directly.
py -3.11 -m venv .venv
.\.venv\Scripts\python.exe -m pip install rdkit==2025.3.1 pubchempy==1.0.5
.\.venv\Scripts\python.exe -m pip freeze
Official sourceExpected result
Packages install; record the resolved versions.
5. Save the cleaning script
All platforms
Save as clean_chemical_data.py. Review the output schema and enable --pubchem only when intended.
"""Preserve original CSV rows while adding RDKit identifiers and optional CIDs.
Sources: https://www.rdkit.org/docs/GettingStartedInPython.html
https://docs.pubchempy.org/en/latest/guide/searching.html
Source-reviewed example; not executed. No salt removal or charge normalization.
"""
import argparse
import csv
import json
import time
from pathlib import Path
from rdkit import Chem, rdBase
def main():
parser = argparse.ArgumentParser()
parser.add_argument("input", type=Path)
parser.add_argument("output", type=Path)
parser.add_argument("--pubchem", action="store_true")
args = parser.parse_args()
added = ["source_row", "canonical_isomeric_smiles", "validation_status",
"duplicate_of_row", "rdkit_version", "pubchem_cids", "lookup_status", "lookup_error"]
with args.input.open(encoding="utf-8-sig", newline="") as source:
reader = csv.DictReader(source)
fields = reader.fieldnames or []
if "smiles" not in fields or len(fields) != len(set(fields)) or set(fields).intersection(added):
parser.error("Require a unique smiles column and no collisions with output columns.")
seen = {}
lookup_cache = {}
# Exclusive creation prevents overwriting either source or existing output.
with args.output.open("x", encoding="utf-8", newline="") as target:
writer = csv.DictWriter(target, fieldnames=fields + added)
writer.writeheader()
for row_number, row in enumerate(reader, start=2):
if None in row:
parser.error("Malformed CSV row; partial output was retained for inspection.")
result = dict(row, source_row=row_number, canonical_isomeric_smiles="",
validation_status="", duplicate_of_row="", rdkit_version=rdBase.rdkitVersion,
pubchem_cids="[]", lookup_status="not_requested", lookup_error="")
smiles = (row.get("smiles") or "").strip()
molecule = Chem.MolFromSmiles(smiles) if smiles else None
if molecule is None:
result["validation_status"] = "empty" if not smiles else "invalid"
else:
canonical = Chem.MolToSmiles(molecule, canonical=True, isomericSmiles=True)
result["canonical_isomeric_smiles"] = canonical
result["validation_status"] = "valid"
if canonical in seen:
result["duplicate_of_row"] = seen[canonical]
else:
seen[canonical] = row_number
if args.pubchem:
if canonical not in lookup_cache:
import pubchempy as pcp
try:
compounds = pcp.get_compounds(canonical, "smiles")
cids = sorted({item.cid for item in compounds if item.cid is not None})
lookup_cache[canonical] = (json.dumps(cids), "matched" if cids else "not_found", "")
except (pcp.PubChemPyError, OSError, ValueError) as error:
lookup_cache[canonical] = ("[]", "error", str(error))
time.sleep(0.3)
result["pubchem_cids"], result["lookup_status"], result["lookup_error"] = lookup_cache[canonical]
writer.writerow(result)
if __name__ == "__main__":
main()
Official sourceExpected result
The script is saved beside .venv; no input rows have been changed.
6. Save the acceptance sample
All platforms
Save as input.csv. It deliberately contains duplicates, malformed SMILES, an empty value and a salt.
id,smiles
ethanol-a,CCO
ethanol-b,OCC
invalid,not-a-smiles
empty,
salt,CC(=O)[O-].[Na+]
Official sourceExpected result
Five data rows are present.
7. Run local cleaning
All platforms
Use the Windows ..venv\Scripts\python.exe interpreter instead on Windows. Do not add --pubchem for this deterministic baseline.
.venv/bin/python clean_chemical_data.py input.csv cleaned.csv
Official sourceExpected result
Five rows retained: three valid, one invalid, one empty. ethanol-b references source row 2 as a duplicate; the salt retains a dot-separated sodium cation and acetate anion. Every lookup_status is not_requested. Input is unchanged.
8. Review the output using this prompt
All platforms
Use this as a human checklist, or provide only the small public sample to an optional assistant. Cloud upload changes the baseline privacy boundary.
Review the generated cleaned.csv, preserving original id and smiles rows. Count valid, invalid, empty and duplicate rows. Explain that ethanol-a and ethanol-b encode the same molecule, and verify that the salt row keeps both charged fragments. Do not remove rows, neutralize salts or infer missing PubChem CIDs. Base the explanation only on the supplied CSV.
Official sourceExpected result
The explanation matches the actual CSV and does not discard or alter structures.
9. Optionally enrich with PubChem
All platforms
Use a different output filename. Record retrieval time and inspect lookup_status/lookup_error; multiple CIDs are retained. Do not call external services on confidential data.
.venv/bin/python clean_chemical_data.py input.csv enriched.csv --pubchem
Official sourceExpected result
Original row count is unchanged; valid rows receive actual CID lists or explicit lookup errors.