Most writing on edge AI deployment strategies asks one question: edge or cloud? For a model that only labels images, that question is fine. For a model whose output moves a valve, a pump or a feed rate, it is the wrong unit of decision. The strategy is not where the model lives. It is where each decision lives, and what happens to that decision when the network, the model or the hardware misbehaves.
This is how we make that call at Equation Labs, where our learned controllers act on live biological and thermochemical processes, and where the numbers below come from measurement rather than a vendor sheet.
A strategy is a placement of decisions, not a choice of box
The most useful framing in the current literature comes from a gateway maker, which is telling. Robustel's workload placement guide argues that the question is not which platform runs the whole application but where each individual workload fits best. Capture, inference, reporting and fleet comparison are different jobs with different constraints, and forcing them into one environment creates a false choice.
The control engineering view points the same way. A 2026 comparison of edge and cloud architectures for industrial control names the three failure modes of a cloud-centred loop: latency, unreliability on connection loss, and the cost of moving raw data. Each of those is a property of a specific decision, not of the whole system. So we start by listing decisions, not servers.
The four strategies
There are four placements worth taking seriously. Real systems combine them.
Cloud round trip
The device streams data to a remote endpoint and gets a decision back. The appeal is real: elastic compute, one deployment to manage, fast rollback. The cost is the round trip. IoT Snacks puts a cloud round trip at roughly 50 to 300+ ms with jitter, against about 1 to 20 ms for small models on a device CPU or NPU and 2 to 30 ms on a local gateway.
Geography sets a floor you cannot optimise away. Measured Azure backbone figures cited by TSL Automation show about 20 ms from Central India to South India but roughly 140 ms to West Europe and 198 ms to East US, before the plant's own link and queueing are added. For a plant in Spain or India, the nearest cloud region decides whether a cloud decision is even plausible.
Site gateway
An industrial PC or gateway on the plant network runs inference for many devices. It keeps raw data on site and survives a WAN outage, while giving you one box per site to update instead of one per sensor. It fits perception, anomaly detection and short-horizon optimisation that must keep working locally but does not own an actuator's inner loop.
Inference at the machine
The model runs on the controller or an accelerator beside it. This is the only placement where latency is bounded by your own scheduling rather than someone else's network. iFactory, which sells edge nodes and should be read with that in mind, gives OT control loops a cycle of 10 to 100 ms and cloud round trips of 80 to 400 ms. If those ranges hold for your process, a cloud path is disqualified from the inner loop before any model question is asked.
The price is that the model must fit. That is where edge AI model compression stops being an optimisation and becomes the entry ticket.
Hybrid and split models
Almost every serious system ends up hybrid: infer locally, send events and a small window of context upstream, retrain centrally, ship updates down. A second hybrid pattern matters more for control. The cloud sends policies (setpoints, thresholds, an allowed operating envelope) and the edge enforces them and runs the loop, which keeps the plant safe through a WAN failure.
A split model, early layers on the device and later layers in the cloud, cuts bandwidth for vision and audio but adds integration work, and embeddings that can be inverted become a privacy problem. We rarely use it for control: it puts a network hop back inside the decision.
Decide by consequence, then by latency
Latency is the number people reach for first. Consequence should come first. Hyperion Consulting's decision architecture separates three planes: a deterministic controller plane for interlocks and closed loops, a local inference plane for perception and short-horizon work, and a fleet plane for training, registries and cross-site analytics. Its design rule is the one we apply: the edge does not silently acquire authority that belongs to a deterministic controller or a human.
So the first question for any decision is: what does a late, wrong or missing output cost? If the answer is damage or injury, the final say sits at the machine, whatever the model's accuracy in the cloud. Only then ask about milliseconds, and ask about the tail, not the mean. A path that averages 100 ms with 500 ms spikes breaks a loop that needs a predictable 50 ms, even though the average looks fine.
Bandwidth, sovereignty and the offline state
Bandwidth often decides the architecture before model size does. iFactory estimates that one line with 200 sensors produces 15 to 40 GB of raw data a day. TSL works through ten 1080p cameras at an assumed 4 Mbps each: about 13 TB a month, with Azure egress from Asia priced at $0.12 per GB after the first 100 GB. Local inference turns that stream into kilobytes of results.
Sovereignty is the quieter driver. Process parameters and recipes are the core IP of many plants, and air-gapped OT networks in energy and pharmaceuticals may forbid the cloud path outright.
Then there is the state everyone forgets. Hyperion treats connectivity as a state, not a boolean: online, intermittent, deliberately offline, degraded, and recovering. A strategy that says "works offline" is incomplete until storage limits, restart while offline, queued commands and reconnection with conflicting state have each been specified and tested. Those tests belong in the specification before the hardware is chosen.
What we run, and why it is split the way it is
Our own control stack is a concrete case. Our published figures are an end-to-end edge latency of 285 ms, cut from 1.2 s, and a safety layer that answers in under 2 ms, validated on a 500 L pilot basin with zero safety violations across 400 simulated years.
Those two numbers are the strategy. The learned model runs at the edge, and the cut from 1.2 s to 285 ms was won on the target hardware, where no network hop can eat into the budget. The safety layer runs beside the actuator, apart from the model, because it has to hold even when the model is wrong or late. It is built on control barrier functions, which change rarely and under the strictest review. Training, long-horizon analysis and the simulation campaigns behind the 400 simulated years run centrally, where compute is elastic and nothing is waiting on them.
Three planes, three cadences. The model can iterate quickly precisely because the layer beneath it bounds what a bad release can do.
The cost the edge adds
Edge placement is not free, and the honest pages say so. Robustel's example is blunt: one cloud deployment becomes 500 when the same function is spread across 500 gateways, each needing version control, resource monitoring, updates and recovery. The best architecture minimises total operational burden, not one metric.
This is why a strategy is only as good as the release process behind it. Choosing the machine as the inference target commits you to a repeatable path from trained model to validated, versioned, reversible release on that hardware. We set out that path stage by stage in our edge AI deployment pipeline. If you are earlier than that and need the ground floor first, start with what edge AI deployment means and how it differs in edge computing vs edge AI.
A decision sheet
| Decision | Default placement | Move it only if |
|---|---|---|
| Safety interlock, protective stop | Controller at the machine | Never to the cloud |
| Closed-loop control output | Machine or real-time compute | Cloud may set bounded targets, never the inner loop |
| Perception, anomaly scoring | Gateway or machine | Latency and consequence tolerate a round trip |
| Operator assistance, reporting | Gateway or cloud | Data cannot leave the site |
| Training, fleet analytics | Cloud or central compute | Site is air-gapped |
Read top to bottom, the sheet is a ladder of consequence. The higher a decision sits, the closer it runs to the process and the slower it is allowed to change.
FAQ
What are the main edge AI deployment strategies?
Four: cloud inference with a round trip, inference on a site gateway, inference on the device or controller, and hybrid designs that combine them. Most production systems are hybrid, with local inference and central training.
Can a cloud model control an industrial process?
It can supervise one. The cloud can send setpoints and an allowed envelope, but the inner loop and the safety interlocks should run locally, so the plant stays safe when the link drops.
When should inference stay in the cloud?
When the decision tolerates hundreds of milliseconds and jitter, when the model is large and changes weekly without reliable over-the-air updates, or for training and multi-site analytics where elastic compute matters more than proximity.
Does edge deployment make model updates harder?
Yes. Every distributed instance needs versioning, monitoring, staged rollout and a tested rollback. That cost is worth paying where latency, bandwidth or consequence demand it, and not elsewhere.

