An AI workflow for chemical engineering should be evaluated as a chain of decisions, not just a predictor with a favourable error score. A process model, an optimisation strategy and a controller answer different questions. This guide maps source-described tool roles, then proposes an evaluation plan. It does not report first-hand scientific testing, plant deployment or controller performance.
Describe the process and its operating envelope
Start with a process specification before selecting software. Identify the equipment boundary, measured variables, manipulated variables, disturbances, units and sampling intervals. State whether the intended output is a monitoring signal, an advisory recommendation or a control decision.
Create a data inventory that includes timestamps, sensor quality flags, laboratory measurement delays, feedstock changes, maintenance events and control-policy changes. Record which operating regimes the data actually cover. A fitted relationship may not remain useful after a change in feed composition or instrumentation.
Input: a process description and traceable measurements. Output: a variable dictionary, regime map and intended-use statement.
Ask: Which variables are available at decision time? Which are measured only later? Does the proposed evaluation cover the conditions in which the result would be used?
Assign each tool a distinct role
The linked sources describe complementary capabilities, not interchangeable products or a verified integrated stack:
- IDAES provides a process systems engineering framework for simulation-based design, analysis and optimisation, with an emphasis on advanced energy systems. Relevant examples and model assumptions must be checked for the selected process; the excerpts do not establish a universal input or output schema.
- SysIdentPy supports nonlinear dynamic-system identification using NARMAX-family models. Its prepared reference describes training input/output arrays, configurable lags, structure selection, parameter estimation, predictions and residual-correlation utilities.
- do-mpc supports nonlinear and economic model predictive control, robust multi-stage control and moving horizon estimation. Its architecture separates simulation, estimation and control, with support for differential algebraic models.
- Summit focuses on iterative reaction optimisation. It provides optimisation strategies and mechanistic or data-driven reaction benchmarks for simulation-based evaluation.
Use this role map to select a component for a defined task. Connecting an identified model to a controller, or transferring reaction-optimisation results into a process simulation, would be a proposed integration requiring interface checks—not an integration demonstrated by the available evidence.
Specify inputs, outputs and objectives
Write a small input/output contract for each evaluation stage. For a dynamic predictor, include input and output histories, lag definitions, preprocessing rules and the forecast horizon. For a process simulation, document parameters, boundary conditions and the quantities to inspect. For an optimisation study, define the search domain, measured objectives and experiment budget. For control evaluation, specify state estimates, references, manipulated variables and timing requirements.
Define success before fitting or tuning. Prediction error, experiment efficiency and control performance require different measures. Include a simple baseline under the same data and operating assumptions, rather than comparing incompatible demonstrations.
Ask: Is the goal to explain observed behaviour, predict a trajectory, recommend the next experiment or select a control action? If several objectives compete, how will trade-offs be reported? Avoid hiding an unacceptable outcome inside an improved weighted score.
Preserve physical and operational constraints
Keep constraints separate from the objective. Express them in the same units and time basis as the model, and distinguish equipment limits, quality requirements, actuator bounds and limits on the rate of change.
Before optimisation, define what counts as a violation, what invalidates a recommendation and what happens when a calculation fails. Check whether a constraint applies to every predicted step or only an endpoint. Decide how uncertain parameters and unavailable measurements will be handled.
Laboratory recommendations should remain separate from commands sent to equipment. A better surrogate objective is not evidence that a plant change is feasible or safe. Likewise, a simulated controller respecting model constraints does not establish that those constraints capture all operational hazards.
Evaluate across time and operating regimes
Partition data to reflect intended use. For forecasting, reserve later periods and keep future observations out of preprocessing, feature selection and parameter tuning. Where relevant, hold out whole campaigns or operating regimes rather than randomly mixing adjacent observations.
For dynamic models, compare trajectories and residual structure as well as aggregate errors. Separate one-step prediction from longer sequential prediction: their information requirements differ. Inspect transitions, missing sensors, delayed measurements and conditions outside the training domain.
For optimisation, propose repeated comparisons under a fixed evaluation budget and report infeasible candidates as well as objective values. Summit's simulated benchmarks offer a setting for such studies, but results on them would not establish performance on a user's chemistry. Its quick-start prose names Nelder-Mead while the code instantiates SOBO; resolve that discrepancy before adapting the example.
Worked planning example: an advisory reactor study
Hypothetical example—not a tested workflow: a team wants an offline advisory model for reactor-temperature behaviour after feed changes. It proposes using minute-sampled feed, coolant and temperature histories, with predictions over the next ten minutes. These timings are illustrative planning choices, not recommended operating settings.
- Define the reactor boundary and confirm which measurements are available when an advisory calculation starts. Treat delayed product-quality measurements separately.
- Reserve the final campaign for evaluation and an earlier campaign for tuning. Fit preprocessing only on training data, and label maintenance and feed-change periods.
- Evaluate a SysIdentPy dynamic model against a simple baseline. Inspect temperature trajectories, regime-specific errors and residual correlations; do not rely only on average error.
- If model behaviour is acceptable under predefined criteria, consider a separate do-mpc simulation study. Verify model translation, state definitions and timing first. There is no assumed plug-and-play connection.
- Include sensor-loss and parameter-mismatch scenarios. Use process-owner-approved limits, not values inferred from historical extrema.
The planned deliverable is an offline report containing prediction errors, constraint checks, rejected cases and unresolved assumptions. IDAES could be considered for a separate physics-based modelling study where an appropriate model is available. Summit would address a different question—iterative reaction-condition optimisation—not automatically replace the dynamic predictor or controller.
Keep deployment accountable: final checklist
The available evidence does not establish plant-specific accuracy, real-time suitability, direct hardware integration or compatibility among these tools. A simulation or fitted model alone also does not establish an operational digital twin.
Before advancing beyond offline evaluation:
- Record model versions, data windows, parameters and preprocessing settings.
- Check interface units, timing and failure behaviour.
- Document acceptance criteria and results by operating regime.
- Define an observable fallback for invalid inputs or out-of-domain conditions.
- Require independent approval for operational-control changes.
- Repeat evaluation after process, feedstock or instrumentation changes.
Plan reproducible evaluations and review local deployment requirements to turn the study into a traceable evaluation package.