Ask whether AI in chemical process control works and you can collect a confident yes and a confident no in the same afternoon, both of them correct. The disagreement is almost never about the method. It is about where in the control hierarchy the model was allowed to sit. A model that infers a laboratory number, a model that raises an alarm, a model that writes a setpoint and a model that moves a valve are four different engineering problems carrying four different burdens of proof, and a great many published results belong to a lower layer than their abstract suggests.
Sorted by layer, the evidence stops contradicting itself. It also becomes usable, because the layer you are proposing tells you exactly which number you are going to be asked for.
The layer decides the question
A plant's control stack is layered deliberately. Field instruments feed a distributed control system running regulatory PID loops. A safety instrumented system sits beside it with interlocks and trips that nothing else may override. Advanced process control or MPC sits above, moving the setpoints those loops track. Optimisation, scheduling and planning sit above that.
Going down the stack, the response deadline shortens and the consequence of being wrong becomes more physical. That is the whole difficulty with putting a learned model into it. As The Chemical Engineer states the tension, modern AI is probabilistic and predicts the most likely outcome from its training data, while process control must be deterministic: a safety shutdown, an interlock or a PID loop demands certainty, and a probabilistic guess is unacceptable when the thing being managed is high-pressure hydrocarbons or an exothermic reaction. The same piece names the second constraint, that many industrial control systems are air-gapped and that security, functional safety and regulatory requirements all limit cloud-based AI.
Neither constraint rules AI out of a chemical plant. Together they decide which layer is worth arguing about.
Layer one: the model measures and does not act
Soft sensing is where most working chemical-industry AI actually lives. A soft sensor predicts an infrequently measured property, usually a laboratory quality result, from the temperature, pressure and flow measurements a plant already takes continuously.
The unglamorous detail is the useful part. In a soft sensor built for the vacuum distillation unit of an Asian refinery, the data set began with 40 physical sensors and only 31 survived into the model. Four were rejected for high noise and large amounts of missing data, and five more were lost because coking on the sixth side draw had corrupted the instruments around it. The target was an ASTM-D2887 boiling curve measured in a laboratory, and the reason to model it at all is turnaround: sample collection, laboratory analysis, reporting. A soft sensor can supply the same information as often as every minute.
Note where the engineering went. Not into the architecture, into deciding which instruments could be believed. Note also the failure mode: a wrong number on a screen, read by a person with the context to distrust it. That is the mildest failure available anywhere in this stack, and it stops being mild the moment the number is consumed by a controller instead of an operator.
Layer two: the model raises an alarm
Fault detection is the most published layer and the most misread, because the metric that dominates its papers is the one that matters least.
On the Tennessee Eastman Process, the standard chemical-process fault benchmark, a recent seven-architecture comparison put LSTM-FCN at 99.37 percent accuracy, a CNN-Transformer at 99.20, a plain LSTM at 99.14 and an optimised XGBoost baseline at 93.91. Read the headline and every temporal deep model looks deployable.
The same study then tested on independently generated data, and that is the result to carry. A convolutional autoencoder that scored 99.39 percent fell to 76.04 percent, a loss of 23.35 points. The LSTM autoencoder gave up 3.02 points over the same shift. The models that degraded worst were the ones with calibrated components, whose parameters had quietly fitted themselves to dataset-specific statistics. Three fault types, 3, 9 and 15, had already been excluded from the benchmark as too subtle to detect reliably.
The numbers a control room asks for
Accuracy is not what a plant buys. Detection delay and false alarms are. A process-guided framework reporting both achieved an average false alarm rate of 0.97 percent with an average detection delay of 5.6 samples after fault onset, and offered operators a selective-risk option to defer the noisiest 10 to 20 percent of cases rather than alarm on them.
There is a third number under both. A physics-guided residual and calibrated CRNN study makes the point that most Tennessee Eastman papers report strong ROC curves while leaving alarm thresholds ad hoc and hard to defend in a control room. Its contribution is probability calibration, an expected calibration error pulled down to roughly 0.03, which is what lets a threshold be set against a stated false alarm budget instead of by taste. Its early-warning governance improved the Numenta Anomaly Benchmark score by 17 percent over an isolation forest baseline.
A model at this layer is arguing for a place in the alarm system, so it should be assessed the way an alarm is: nuisance rate first.
Layer three: the model writes a setpoint
This is where the credible deployments cluster, because the regulatory layer and its fallback both stay exactly where they are. The learned component proposes a target; a controller that has been trusted for twenty years still holds it.
The sharpest evidence is a four-way ladder run on Skogestad's Column A distillation benchmark with identical level closure, scenarios and seeds: PID alone, linear MPC, a learned supervisor, and the same supervisor behind a safety gate. On off-nominal target acquisition the learned supervisor beat Pareto-tuned linear MPC hard, an IAE ratio of 0.361 at the upper confidence bound. On steady-state disturbance rejection, same column, same seeds, the result inverted by 10.18 at the point estimate. That is not an unresolved argument. It is a clean split by regime, and which method wins on your own process is predictable before you build the simulator.
The other move at this layer is to stop treating the model as a black box over the data. A deep network built on process mechanism knowledge, gated recurrent units with attention arranged in a distributed structure with residual connections, was used inside predictive control of a distillation process precisely because purely mechanistic and purely data-driven modelling each fail differently. Interpretability is not decoration here. Someone has to sign the setpoint off.
Layer four: the valve, and the single case that reached it
There is one documented field result of AI holding a chemical plant directly, and it is worth reading closely rather than citing as a trend.
In 2022, Yokogawa and JSR ran an AI controller autonomously on a chemical plant for 35 consecutive days. The controlled unit was a distillation column, and the task was one that PID control and APC could not handle: holding product quality and column liquid level while making maximum use of waste heat, under rain and snow that swung the ambient temperature. Previously that combination required an experienced operator to suspend automatic control and work the valves by hand. Over the test the products met their standards and were shipped, and the off-spec losses that normally accompany manual intervention did not occur.
Now the part a press release does not lead with. The FKDPP reinforcement learning protocol behind it was developed with the Nara Institute of Science and Technology in 2018, validated on a control training system in 2019, then on a simulator recreating an entire plant in April 2020, before it was allowed near the real column. The field test ran under a Japanese Ministry of Economy, Trade and Industry advanced industrial safety subsidy. So the honest summary is one column, 35 days, four years of staged validation and a state safety programme behind it. That is the price of layer four today, and it is the reason the layer is nearly empty.
What has to be true before a model touches an actuator
Four conditions, each one a question the layers above have already answered somewhere.
The worst-case inference time fits inside the sampling period. Not the mean. In a benchmark chemical process at UCLA, a learned controller ran at 0.644 ms average and 67.5 ms worst case while a short-horizon linear MPC took 15.247 ms mean and 112.516 ms worst case, both comfortably inside a one second budget. The mean hides the step that would have missed the deadline.
A safety mechanism sits outside the learned policy. In every study above that claims a guarantee, the guarantee lives in a component the policy cannot rewrite: a gate, a backup controller, an enforcer. The most portable version is a filter wrapped around a policy you already have, which is also the pattern the gated configuration on Column A used.
Inference runs at the plant. The control network is air-gapped, so a cloud round trip is not merely slow, it is unavailable. That constraint usually turns into a model compression problem with known answers rather than a networking one.
There is a retraining path, and it is costed. Changing a learned policy's behaviour means changing its objective and retraining, where changing an MPC's behaviour means editing a cost or a constraint. On a plant whose feed, spec or tariff moves several times a year that difference is the maintenance budget, and it is why the gap between simulator and plant has to be measured as a quantity rather than noted as a caveat. It also puts a hard requirement on the training environment, since a simulator you intend to certify against is a different artefact from one built to explore ideas. The same four conditions apply outside the chemical vertical too, with different economics on discrete and utility plant.
Which layer our own stack sits on
Ours sits at layer three with a layer-four safety component, on a biological cultivation process rather than a chemical plant. The published figures are end-to-end edge latency of 285 ms, down from 1.2 s, a safety layer answering in under 2 ms, and a growth model at R-squared above 0.95, validated on a 500 L pilot basin with zero safety violations across 400 simulated years.
Those numbers answer two of the four conditions above and nothing else. The 285 ms is an answer to the deadline question, and the 400 simulated years is an answer to the validation question. They do not say a learned controller outperformed a well-tuned conventional one on that process, because that is not the comparison they were collected to settle, and a distillation column is not a cultivation basin. What transfers between the two is the layering and the order of proof, not the figures.
FAQ
Can AI replace PID controllers in a chemical plant?
Essentially never as a replacement. Published deployments leave the regulatory PID layer and the safety instrumented system in place and put the learned component above them. Even the Yokogawa and JSR field test was aimed at conditions PID and APC could not handle at all, rather than at loops that were already working.
Has AI ever actually run a chemical plant on its own?
Once, on record. In 2022 an AI controller ran a JSR distillation column autonomously for 35 consecutive days and the product met specification and was shipped. It was one column over 35 days, following four years of staged validation on a training system and a whole-plant simulator.
What is a soft sensor in chemical process control?
A model that predicts an infrequently measured property, typically a laboratory quality result such as a boiling curve, from the temperature, pressure and flow measurements taken continuously by physical instruments. It replaces a laboratory turnaround measured in hours with a prediction available as often as every minute.
Why do fault detection models with 99 percent accuracy fail on a real plant?
Because the accuracy was measured on the distribution the model was fitted to. On independently generated Tennessee Eastman data one convolutional autoencoder fell from 99.39 to 76.04 percent, and the components with calibrated parameters degraded the hardest.
Does AI process control need a cloud connection?
Not for anything inside a control loop. Many industrial control systems are air-gapped by design, and functional safety and regulatory requirements constrain cloud AI further, so the model has to run on local hardware at the process.

