How it works
A model is first trained with plenty of computing power, usually in the cloud. It is then shrunk so it fits a small device, using techniques such as quantisation (storing numbers with fewer bits) and pruning (removing parts that matter little). The device runs it on an efficient chip, often a neural processing unit, and acts on the result at once.
Why run AI at the edge
Speed: a decision made on the device does not wait for a round trip over the network. Privacy: raw video or sensor data can stay where it was captured. Cost: sending less data saves bandwidth. Resilience: the system keeps working when the connection is slow or down, which matters on a factory floor, a farm or a ship.
Trade-offs
A small device has less memory and power than a data centre, so edge models are smaller and sometimes less capable. Fleets of devices also need updating and securing, which takes planning. Many systems use both: quick decisions at the edge, heavier analysis and retraining in the cloud.
Where it is used
- Cameras Detecting a fault, a vehicle or a person on site without streaming the video away.
- Machines Spotting unusual vibration or sound and warning before a breakdown.
- Phones and wearables Recognising speech, photos or a heartbeat on the device itself.
- Vehicles Making decisions in milliseconds where waiting for a network is not an option.
Edge AI compared with cloud AI
Cloud AI has near-unlimited computing power and the largest models, but it needs a connection, adds delay and moves data off the site. Edge AI gives up size in return for speed, privacy and independence.
Most real systems use both: a small model decides on the spot, and the cloud retrains it and handles the heavy analysis. Choose the edge when the answer is needed in milliseconds, when the connection is unreliable, or when the data should not leave the site.
Key terms
- Inference
- Running a trained model on new data to get a result.
- Quantisation
- Storing a model’s numbers with fewer bits so it is smaller and faster.
- NPU
- A neural processing unit: a block of a chip built to run neural networks efficiently.
- Latency
- The delay between an input and the system’s response.
Common questions
What is the difference between edge AI and cloud AI?
- Cloud AI runs on remote servers and needs a connection. Edge AI runs on or near the device that produces the data.
Does edge AI need the internet?
- It can run without it. Many setups still connect now and then to send summaries, receive updated models or report faults.
What hardware runs edge AI?
- Anything from a microcontroller with a small accelerator to a rugged computer with a dedicated neural processor or graphics chip, depending on how large the model is.
Sources and further reading
Independent pages we checked while writing this guide. They are not Reinnder products, and Reinnder is not affiliated with them.
Reinnder’s angle
How Reinnder looks at aI on the device
Reinnder Silicon is the planned home for chips and edge computers designed to run models on the machine itself.
This guide explains the technology in general terms. It is not advice, and it does not describe a Reinnder product on sale.