Zero shot sim to real transfer is the practice of taking a policy trained entirely in simulation and running it on the physical system with no further training on real data. The standard statement of it is deliberately absolute: all learning and policy development happens in simulation, and the resulting controller is deployed directly. It stands in contrast to the methods that spend real data to close the gap, fine-tuning, residual learning, and system identification performed after the simulator has already trained something.
Read that carefully and it is not a description of a technique. It is a description of a budget. Zero shot says nothing about how well the controller works. It says only that nobody paid for a calibration campaign on the real machine, which is a claim about your evidence, and a claim you have to be entitled to make before you put a controller on a process that can be damaged.
The term has two meanings and they skip opposite things
In robotics, zero shot means no real data was used to adapt the policy. The simulator is the expensive artefact, and it is assumed to exist.
In building control the word points the other way. The PEARL work on zero shot building control uses it for fitting a policy online with no simulator and no historical data at all, given only a short commissioning period, and reports cumulative emissions up to 31.46 percent below a rule based controller. Its motivation is that building an accurate simulator in the standard software can take an expert months and is impossible without knowledge of the building's topology and thermal parameters. There, the simulator is the thing you were allowed to skip.
Both usages are established and neither is wrong. But a reader comparing two zero shot results is often comparing claims about two different skipped artefacts, and the first thing worth writing down in any transfer plan is which one you mean.
The published record understates the calibration that already happened
The most useful paper in this literature is the one that went looking for the best method and found a problem with the reporting instead. Valassakis, Ding and Johns benchmarked zero shot transfer for dynamics across real tasks and state the finding without hedging: many works do not present a thorough evaluation in the real world, or underplay the significant engineering effort and task-specific fine tuning that is required to achieve the published results. Others bypassed the real problem entirely by validating in simplified simulated environments, presenting sim-to-sim results rather than sim-to-real ones.
The simplest randomisation matched the complex ones
Their benchmark result is the practically important half. Simply injecting random forces into the simulation performed at least as well as randomising the full set of dynamics parameters, and as well as adapting a policy online with a recurrent architecture. Their conclusion is that many of the more complex recent methods do not scale to real-world tasks without significant task-specific tuning, which defeats the purpose of zero shot transfer.
That is worth holding onto when a vendor or a paper offers you a sophisticated adaptation scheme. If the scheme needs per-task tuning against the real system, the tuning is the calibration campaign, and the zero shot label has quietly moved rather than disappeared.
The baseline was fitted to real data before any randomisation began
The same paper describes how it set its baseline parameters, and this is the detail that decides the question for an industrial process. Kinematics came from the robot's description file. Geometry came from physical measurement. Friction, notoriously hard to measure, came from educated guesses at typical material values. And the arm's dynamics parameters, controller gains, damping ratios and joint friction, were optimised by differential evolution against the simulated robot's response to a real control signal.
That last step is system identification. It happened before the randomisation, against real data, and it does not count against the zero shot claim because the claim is scoped to policy training. On a robot arm that accounting is harmless, because collecting a step response is cheap. On a gasifier or a cultivation basin the response campaign is the expensive part, and if you inherit the robotics recipe you inherit an unbudgeted line item that dwarfs the training run.
What the bounds say you must have
There is theory here, and it prices the question rather than answering it. Work on the statistical guarantees for offline domain randomization starts from the known result that uniform randomization does bound the sim-to-real gap, then points out how badly the bound scales. For a finite, separated class of candidate simulators the performance gap between the optimal policy on the true system and the domain-randomized policy is O(M cubed log(MH)), with M the number of candidate simulators and H the horizon. Without the separation condition it is O(square root of M cubed H log(MH)).
The alternative is to fit the randomization distribution to an offline dataset from the real system before training, which the authors formalise as maximum likelihood estimation over a parametric simulator family. They prove it is weakly consistent under regularity, positivity and identifiability assumptions, and strongly consistent with one extra Lipschitz continuity assumption, meaning the fitted distribution converges to the true dynamics as the dataset grows. The bounds improve to O(M squared log(MH)) and O(square root of MH log(MH)) respectively, a factor of O(M).
Treat M as a bill. Every parameter you decline to identify stays in the candidate set and is charged at that rate. This is the argument for spending your scarce real data on the distribution rather than on the policy, and it is a separate question from where the randomization ranges come from once you have decided to fit them.
One caution from the same work: their E-DROPO variant adds an entropy bonus specifically to prevent variance collapse, yielding broader randomization and more robust zero-shot transfer in practice. Fitting the distribution too tightly to the offline record is its own failure mode, not a virtue.
One controller, two regimes, two outcomes
The cleanest evidence about zero shot on a real process is not in the robotics literature at all. A field evaluation of an RL HVAC controller trained the policy entirely in a data-driven digital twin of an office building, then deployed it across two air handling units under two scenarios.
Under static thermal comfort limits, the transfer succeeded. The controller consistently maintained thermal comfort and achieved energy use comparable to the existing building management system. Under dynamic comfort limits, performance degraded, and the authors attribute it to reward design vulnerabilities, policy generalization limits and the sim-to-real gap, calling out sensitivity to nonstationary environments.
Same controller. Same simulator. Same building. The zero shot claim held in one operating regime and not in the other, which means transferability is a property of the policy and the regime together, not a property of the method you used. Any go or no-go decision that names only the method is answering the wrong question.
A deployment on a building's thermally activated system is instructive for what it did differently. The Soft Actor-Critic agent was pre-trained on a simplified resistance-capacitance model calibrated with real building data from the same testbed, and the on-site implementation included a fail-safe mechanism. It also benchmarked the learned controller against the incumbent rule based controller during real operation, and the authors name establishing that benchmark as a primary contribution in its own right. Calibrated surrogate, independent fail-safe, honest comparator: three choices that are easy to read as implementation housekeeping and are in fact the reason the deployment was defensible.
Four conditions before a controller goes on without calibration
This is the rule we apply before a policy is allowed onto a live physical or biological process with no real-system campaign behind it.
The regime is the one you trained in, and it is stationary. The HVAC result is the warning. Nonstationary setpoints, a changing feedstock or a seasonal drift put you outside the distribution the policy was optimised over, and nothing about a zero shot method covers that.
The dominant gap is parametric rather than structural. Randomization covers mismatch in quantities your simulator has a parameter for. A phenomenon the model does not represent is untouched by it, which is the harder half of the reality gap on a live process and the reason a state model that resolves the structure matters more than a wider distribution.
You hold offline records good enough to fit the distribution to. Without them you are on the uniform randomization bound, paying the O(M cubed) rate on every parameter you guessed. Historian data you already own is the cheapest way off that curve, and it is also what makes a digital twin good enough to certify against worth building.
Something independent holds the constraint at runtime. A learned policy that transfers imperfectly should cost you yield, not equipment. That requires a layer outside the policy, which is what a control barrier function makes a hard constraint rather than a trained tendency, and it is the condition that changes a failed transfer from an incident into a data point. It is also the honest comparison point against how a learned controller compares with model predictive control, where the constraint handling is inside the optimiser instead.
Fail any of the four and the correct description of what you are doing is few shot transfer with the shots left uncounted.
What our own numbers do and do not prove
Our control stack carries the figures we quote for this: growth-model accuracy at R-squared above 0.95, end-to-end edge latency of 285 ms reduced from 1.2 seconds, and validation on a 500 L pilot basin with zero safety violations across 400 simulated years.
We say the same thing about that last number wherever it appears, because it is the one most easily misread as a transfer guarantee. Four hundred simulated years is coverage of the state space under an assumed distribution. It is not elapsed operating time, and it is not evidence that the next deployment will be zero shot. A simulator that agrees with itself for four hundred years has demonstrated the width of its own assumptions and nothing else. What made that controller safe to run was the layer underneath it that did not depend on the simulation being right.
FAQ
Is zero shot sim to real transfer the same as domain adaptation?
No, it is the case where domain adaptation was not used. Domain adaptation covers the strategies that spend some real-world data to close the gap: fine-tuning, residual learning, and system identification carried out after simulation training. Zero shot is the stricter regime in which all learning happens in simulation and the resulting controller is deployed directly, so success is measured on the first real attempt.
How do you prove a zero shot transfer actually worked?
Against a benchmark controller on the real system, never against the simulator. The building deployments cited above compare the learned controller with the incumbent rule based or building management system controller during live operation, and one of them names establishing that benchmarking procedure as a primary research contribution. Simulation return is a hypothesis about the plant, not a measurement of it.
Does domain randomization guarantee zero shot transfer?
It bounds the gap under conditions, and the bound is expensive. Uniform randomization scales cubically in the number of candidate simulators in the finite separated case, which is another way of saying that a policy trained across a large family of guessed dynamics carries a weak guarantee. Fitting that family to offline data first tightens the bound by a factor of O(M), so the guarantee is real but priced by how much you refused to identify.
Can you do zero shot control without a simulator at all?
In building control, that is exactly what the term is used for: fitting a policy online with no simulator and no historical data, using a short commissioning period, with reported emissions up to 31.46 percent below a rule based controller. It is a different claim from the robotics one and the two should not be compared as if they were the same result.
Which transfer method should I start with on a process controller?
The simplest one that covers the disturbance class you care about. The only benchmark to test this systematically on real hardware found that injecting random forces into the simulation performed at least as well as randomising the full dynamics parameter set or adapting online with a recurrent policy, while being much easier to implement and interpret. Start there, and add complexity only when a measured failure justifies it.

