Skip to content
Official

Chemical Data Cleaning Stack

Validate and canonicalize SMILES with RDKit without dropping original rows; flag duplicates and optionally enrich valid entries with PubChem identifiers.

Level: Intermediate Cost: Free Privacy: Depends on configuration ~30 min
Start Setup

You'll be able to

  • Preserve original data and label invalid or missing SMILES.
  • Flag canonical duplicates and optionally retain PubChem CIDs.

What you'll build

Scope

Use Python 3.11, RDKit 2025.3.1 and optional PubChemPy 1.0.5. The included script preserves original columns and adds canonical isomeric SMILES, validation state, source row and duplicate reference. It neither deletes invalid rows nor performs automatic salt removal, charge neutralization or tautomer standardization. Canonical SMILES is a representation policy, not proof that every structure is scientifically equivalent.

Privacy and output

Baseline cleaning is local. The --pubchem flag sends valid canonical SMILES to external PubChem, so review dataset confidentiality before enabling it. The script refuses to overwrite an existing output and caches repeated lookups within a run. It preserves multiple CIDs and lookup errors rather than selecting an arbitrary match. PubChem queries may be slow or rate limited; no runtime guarantee is asserted.

SMILES validation and canonical representation

RDKit

Optional PubChem enrichment

PubChemPy

Stack Components

RDKit

SMILES validation and canonical representation · ==2025.3.1

Preserves stereo, charges and disconnected fragments; original input columns are retained.

See upstream licenses and client/service account terms.

View Resource

PubChemPy

Optional PubChem enrichment · ==1.0.5

Only used with --pubchem; sends valid SMILES to PubChem and preserves all returned CIDs.

See upstream licenses and client/service account terms.

View Resource

Compatibility

ClientOSArchitectureVersion requirements
Python macOSAny>= 3.11
Python WindowsAny>= 3.11
Python LinuxAny>= 3.11

Setup & Test

1. Prepare a small CSV and Python

All platforms

Install Python 3.11. Work on a copy of the dataset with a unique smiles column; choose a new output filename. Decide explicitly whether external PubChem enrichment is permitted.

Official source

Expected result

Input copy and unused output path are prepared.

2. Install isolated Python dependencies

macOS

Run in a new working directory. This shell variant applies to macOS/Linux; use the separate Windows step on Windows.

python3.11 -m venv .venv
.venv/bin/python -m pip install rdkit==2025.3.1 pubchempy==1.0.5
.venv/bin/python -m pip freeze
Official source

Expected result

Packages install; record the resolved versions.

3. Install isolated Python dependencies

Linux

Run in a new working directory. This shell variant applies to macOS/Linux; use the separate Windows step on Windows.

python3.11 -m venv .venv
.venv/bin/python -m pip install rdkit==2025.3.1 pubchempy==1.0.5
.venv/bin/python -m pip freeze
Official source

Expected result

Packages install; record the resolved versions.

4. Install Python dependencies on Windows

Windows

Use PowerShell and the virtual-environment interpreter directly.

py -3.11 -m venv .venv
.\.venv\Scripts\python.exe -m pip install rdkit==2025.3.1 pubchempy==1.0.5
.\.venv\Scripts\python.exe -m pip freeze
Official source

Expected result

Packages install; record the resolved versions.

5. Save the cleaning script

All platforms

Save as clean_chemical_data.py. Review the output schema and enable --pubchem only when intended.

"""Preserve original CSV rows while adding RDKit identifiers and optional CIDs.

Sources: https://www.rdkit.org/docs/GettingStartedInPython.html
         https://docs.pubchempy.org/en/latest/guide/searching.html
Source-reviewed example; not executed. No salt removal or charge normalization.
"""
import argparse
import csv
import json
import time
from pathlib import Path

from rdkit import Chem, rdBase


def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("input", type=Path)
    parser.add_argument("output", type=Path)
    parser.add_argument("--pubchem", action="store_true")
    args = parser.parse_args()
    added = ["source_row", "canonical_isomeric_smiles", "validation_status",
             "duplicate_of_row", "rdkit_version", "pubchem_cids", "lookup_status", "lookup_error"]
    with args.input.open(encoding="utf-8-sig", newline="") as source:
        reader = csv.DictReader(source)
        fields = reader.fieldnames or []
        if "smiles" not in fields or len(fields) != len(set(fields)) or set(fields).intersection(added):
            parser.error("Require a unique smiles column and no collisions with output columns.")
        seen = {}
        lookup_cache = {}
        # Exclusive creation prevents overwriting either source or existing output.
        with args.output.open("x", encoding="utf-8", newline="") as target:
            writer = csv.DictWriter(target, fieldnames=fields + added)
            writer.writeheader()
            for row_number, row in enumerate(reader, start=2):
                if None in row:
                    parser.error("Malformed CSV row; partial output was retained for inspection.")
                result = dict(row, source_row=row_number, canonical_isomeric_smiles="",
                              validation_status="", duplicate_of_row="", rdkit_version=rdBase.rdkitVersion,
                              pubchem_cids="[]", lookup_status="not_requested", lookup_error="")
                smiles = (row.get("smiles") or "").strip()
                molecule = Chem.MolFromSmiles(smiles) if smiles else None
                if molecule is None:
                    result["validation_status"] = "empty" if not smiles else "invalid"
                else:
                    canonical = Chem.MolToSmiles(molecule, canonical=True, isomericSmiles=True)
                    result["canonical_isomeric_smiles"] = canonical
                    result["validation_status"] = "valid"
                    if canonical in seen:
                        result["duplicate_of_row"] = seen[canonical]
                    else:
                        seen[canonical] = row_number
                    if args.pubchem:
                        if canonical not in lookup_cache:
                            import pubchempy as pcp
                            try:
                                compounds = pcp.get_compounds(canonical, "smiles")
                                cids = sorted({item.cid for item in compounds if item.cid is not None})
                                lookup_cache[canonical] = (json.dumps(cids), "matched" if cids else "not_found", "")
                            except (pcp.PubChemPyError, OSError, ValueError) as error:
                                lookup_cache[canonical] = ("[]", "error", str(error))
                            time.sleep(0.3)
                        result["pubchem_cids"], result["lookup_status"], result["lookup_error"] = lookup_cache[canonical]
                writer.writerow(result)


if __name__ == "__main__":
    main()
Official source

Expected result

The script is saved beside .venv; no input rows have been changed.

6. Save the acceptance sample

All platforms

Save as input.csv. It deliberately contains duplicates, malformed SMILES, an empty value and a salt.

id,smiles
ethanol-a,CCO
ethanol-b,OCC
invalid,not-a-smiles
empty,
salt,CC(=O)[O-].[Na+]
Official source

Expected result

Five data rows are present.

7. Run local cleaning

All platforms

Use the Windows ..venv\Scripts\python.exe interpreter instead on Windows. Do not add --pubchem for this deterministic baseline.

.venv/bin/python clean_chemical_data.py input.csv cleaned.csv
Official source

Expected result

Five rows retained: three valid, one invalid, one empty. ethanol-b references source row 2 as a duplicate; the salt retains a dot-separated sodium cation and acetate anion. Every lookup_status is not_requested. Input is unchanged.

8. Review the output using this prompt

All platforms

Use this as a human checklist, or provide only the small public sample to an optional assistant. Cloud upload changes the baseline privacy boundary.

Review the generated cleaned.csv, preserving original id and smiles rows. Count valid, invalid, empty and duplicate rows. Explain that ethanol-a and ethanol-b encode the same molecule, and verify that the salt row keeps both charged fragments. Do not remove rows, neutralize salts or infer missing PubChem CIDs. Base the explanation only on the supplied CSV.
Official source

Expected result

The explanation matches the actual CSV and does not discard or alter structures.

9. Optionally enrich with PubChem

All platforms

Use a different output filename. Record retrieval time and inspect lookup_status/lookup_error; multiple CIDs are retained. Do not call external services on confidential data.

.venv/bin/python clean_chemical_data.py input.csv enriched.csv --pubchem
Official source

Expected result

Original row count is unchanged; valid rows receive actual CID lists or explicit lookup errors.

Troubleshooting

  • CSV header error: use a unique smiles column and avoid generated-field collisions.
  • Invalid SMILES: preserve the raw row and correct it from its source.
  • Output already exists: select a new name; do not overwrite the input.
  • PubChem unavailable or rate limited: retain local results and lookup_error, then retry a small subset later.
  • Multiple matches: retain every CID and inspect structures before selecting.
  • Partial output after malformed CSV: inspect and correct input before rerunning with a new output name.
Still not working

Alternatives

Use RDKit alone without PubChem for local validation. Salt/tautomer normalization requires a separately documented scientific policy and a separate output, rather than silently changing this script.