Read four papers on reinforcement learning vs MPC for process control and you get four different answers. On a continuous stirred tank reactor, one comparison finds RL superior to nonlinear MPC and credits the infinite-horizon objective for it. On the Van de Vusse benchmark, NMPC comes in 20 percent lower on integral absolute error than every RL algorithm tested. On a polymerisation reactor, TD3 beats MPC on the reagent molar equivalent under high disturbance while MPC is more robust on reactor temperature, in the same experiment. On an actual methanol distillation column, the learned agent responded more slowly than the model-based method to an unmodelled weather disturbance and the bottom product quality suffered for it.
Those results are usually read as an unsettled argument. They are not. They only conflict if you read each one as a verdict on a method. Sorted by regime instead, they agree almost perfectly, and the agreement is specific enough to predict which side your own process will land on before you spend a year building a simulator to find out.
What each method is actually spending
The mechanism is short, because anyone searching this phrase already half knows it. MPC carries a model, simulates it forward, and solves a constrained optimisation online at every control step. RL moves that cost offline into training and leaves behind a stored policy that is a function evaluation at runtime.
Everything downstream follows from where the compute sits. The survey that classifies the two families puts it as near-orthogonality: MPC is favoured where measurement data is scarce and expensive and the environment admits an optimisation-friendly model, RL where interaction data can be generated in quantity. A review of RL in process control states the industrial half of that: MPC and real-time optimisation are established technologies at the minute and hour decision scale, but they depend on complex models with periodic recalibration, while RL offers adaptive control at low computational cost after training and pays for it with extensive offline learning.
That is the whole trade. The regimes below are just the places where it comes out differently.
Where reinforcement learning wins
The objective is infinite horizon and the controller is not
The CSTR result is the clearest mechanistic win. RL learns a policy against an infinite-horizon return; MPC optimises over a finite horizon and then re-optimises. Where the process rewards a decision whose payoff arrives past the horizon, the horizon is where the performance goes. Lengthening it costs solve time, which is the next regime.
The online optimisation does not fit inside the sampling period
This is the argument that survives contact with a control cabinet. In a benchmark chemical process at UCLA, the RL controller ran at 0.644 ms average and 67.5 ms worst case, comfortably inside a one second sampling budget, while a short-horizon linear MPC took 15.247 ms mean and 112.516 ms worst case. That short horizon was not free either: short-horizon LMPC and a P controller both sat around 4.8 percent higher in closed-loop cost than the long-horizon baseline. The RL policy bought the long-horizon objective at short-horizon compute.
The Van de Vusse study reaches the same place from the other direction. At the process inversion point the RL algorithms and NMPC stabilised the reactor with no significant difference in IAE, and the stated advantage of the learned controllers was purely computational: the action is stored in the policy rather than recomputed by an optimiser every iteration.
The objective is economic and nobody has written it as a tracking cost
MPC needs a cost function. Some plants do not have one they can write down. In a two-stage grinding circuit study by ABB, Boliden and KTH, PPO was pointed at a profit function supplied by the mine operator rather than at a setpoint, and compared against a PID strategy hand tuned over years of plant experience. In some operating cases it controlled the circuit more efficiently. The reason that matters commercially: grinding accounts for 47 percent of the cost per concentrated copper ton, so the objective worth optimising was never particle size on its own.
Where MPC wins, which is the half most comparisons skip
Disturbance rejection, in source after source
The sharpest evidence is a four-way ladder run on Skogestad's Column A distillation benchmark: PID only, linear MPC, a learned supervisor, and the same supervisor behind a safety gate, with identical level closure, scenarios and seeds. On off-nominal target acquisition the learned supervisor beat Pareto-tuned linear MPC hard, an IAE ratio of 0.361 at the upper confidence bound. On steady-state disturbance rejection, on the same 16-point grid, the result inverts by 10.18 at the point estimate. Same plant, same seeds, opposite answers, split cleanly by regime. Two caveats belong with that number: the supervisor there is an LLM writing setpoints every five minutes, not an RL policy, and the headline numbers are single-column and model-conditional. What transfers is the shape of the split, not the magnitude.
The methanol column agrees on real hardware, where the learned method was slower to reject an unpredicted disturbance than model-based estimation and prediction.
Anything outside the training distribution
A policy encodes the disturbances it saw. The PC-Gym benchmark makes this concrete by using an NMPC with perfect state estimation as an oracle rather than an opponent, then testing outside the training distribution. Inside it, the SAC agent had the better mean normalised optimality gap, 76.58 against 138.97 for DDPG. On a step change to 375 K from outside that distribution, the ranking swapped: DDPG at 408.3 against SAC at 525.08, because the aggressive policy that won in training now spiked the reactor temperature. The ranking you measure in the simulator is not a property of the algorithms. It is a property of the disturbance set, and it does not have to survive the plant. This is the same reason sim-to-real gaps have to be measured as a quantity rather than noted as a caveat.
The specification changes on a Tuesday
The methanol distillation study names the operational cost almost nobody benchmarks: changing an RL agent's behaviour means changing the reward and retraining, which is slow, whereas an MPC's behaviour changes as soon as you change its cost or constraints. On a plant whose product spec, feed or tariff moves several times a year, that is not a footnote. It is the maintenance budget.
Five questions that predict your regime
- Is the hard part a constraint or an objective? MPC satisfies constraints because they are inside the problem it solves; any solution is feasible by construction. RL reaches them through penalties, or through a mechanism bolted on outside the policy. A yes here usually ends the discussion before the other four questions get asked, and if it does not, read which guarantee you actually need first.
- Does the solve fit in the control period, worst case rather than mean? Compare your own numbers to the 112.516 ms worst case above against a one second period. Mean solve times hide the step that matters.
- Can you enumerate your disturbances, or only the ones you have already seen? If the second, weight the out-of-distribution result heavily.
- Is the objective a setpoint or an economic quantity nobody has written down? The grinding circuit is the case for the second.
- How often does the specification change? Count retraining cycles, not training cycles.
The mature answer is almost never a replacement
The field's own response to the orthogonality was not to pick a side. It was to combine, which is what the survey above spends its length classifying. The most deployable version for a plant that already runs MPC is to pretrain the RL actor and critic from the MPC's own calculations, so the agent begins by imitating a controller that is already trusted and improves from there, rather than exploring on a live reactor. The Column A study's practical reading is the same shape from the safety side: MPC as the default for local regulation, the learned layer only where re-planning is genuinely warranted, and never as wholesale MPC replacement.
Note what the gated configuration in that study actually is. A trusted simple mechanism checks an untrusted complex controller before anything moves, which is the runtime-assurance pattern, and the same idea as a runtime filter around a policy you already have. The gate compressed a failure mode that would otherwise have run away into a bounded offset. Architecture did that, not training.
What the choice looked like on a live process
Our own control stack sits on a biological cultivation process, and the decision was made on the compute and constraint axes rather than on benchmark IAE. The published figures are end-to-end edge latency of 285 ms, down from 1.2 s, with the safety layer answering in under 2 ms, a growth model at R-squared above 0.95, validated on a 500 L pilot basin with zero safety violations across 400 simulated years.
Read those honestly. They say the stack was fast enough for the loop it sits in and that the safety layer was checked before it governed anything alive. They do not say a learned policy beat an MPC on that process, because that is not the comparison those numbers were collected to settle. The 285 ms figure is an answer to question two on the list above, and the 400 simulated years is an answer to question three.
Running the comparison on your own process
Four things make the experiment worth the time.
Use an NMPC with full state knowledge as an oracle rather than as an opponent, the way PC-Gym does. Reporting an optimality gap against a near-optimal reference makes a bad RL number diagnostic instead of rhetorical.
Test deliberately outside the disturbance distribution, and treat a rank inversion there as the primary result rather than an appendix.
Measure worst-case solve time against the control period, on the hardware that will run it. If the model has to shrink to make that budget, that is a model compression problem with known answers, and it is a better problem than a horizon you had to cut.
Keep the baseline honest. The Column A ladder relay-tuned its PID baseline through a pre-registered shootout specifically so the comparison could not be dismissed as beating a strawman, and its linear MPC still beat that baseline by 0.122 against 0.836 in aggregate IAE. A learned policy that only beats an untuned PID has told you nothing about MPC. Where the plant data arrives irregularly, the model side of this comparison has its own literature, starting with fitting a continuous-time model to irregular plant data. The narrower version of this same argument against the incumbent baseline is worth running too: the same comparison against a PID baseline answers a different question, and so does the quality of the simulator the policy is trained against.
FAQ
Is reinforcement learning better than MPC for process control?
Not as a general statement. On a CSTR one study finds RL superior because it solves an infinite-horizon problem, and on the Van de Vusse benchmark NMPC comes in 20 percent lower on integral absolute error. The pattern that holds across sources is by regime, not by method.
Can reinforcement learning guarantee constraint satisfaction the way MPC does?
Not by itself. In the studies that do claim a guarantee, it sits outside the learned policy: a Lyapunov-based stability enforcer with a backup controller, or a rule-based gate. The UCLA work is explicit that with such an enforcer active, no rigorous claim can be made that the learned controller was the one acting at most sampling times.
How much faster is a trained policy than MPC at runtime?
In the UCLA benchmark, 0.644 ms mean and 67.5 ms worst case for the RL controller against 15.247 ms mean and 112.516 ms worst case for a short-horizon linear MPC, both inside a one second period. Compare worst case to your control period, never the mean.
Why do RL controllers fail on disturbances they were not trained on?
Because the disturbance distribution is encoded in the weights. PC-Gym showed the ranking of two agents inverting between an in-distribution and an out-of-distribution step change, so the better agent in training was the worse one on the plant.
Can you use MPC and reinforcement learning together?
That is where most current research sits. One practical form pretrains the RL agent on the calculations an existing industrial MPC already produces, so it starts by matching MPC performance and improves without a period of unsafe exploration.

