- Stage
- Now · 2026 to 2027
- Area of work
- Reinnder Silicon
- Used in
- Phones and laptops · Devices in the field
Figure · Illustration
Four ways to make a small model capable
-
Train on carefully chosen, high-quality text, so that a small model learns more from less.
-
Teach the small model to imitate the answers of a larger one.
-
Store the numbers with fewer bits, so the model fits in less memory and runs faster.
-
Train a little further on examples from one task, so the model is sharp at that job.
What counts as small#
Size is measured in parameters, the numbers a model learns. The largest models have hundreds of billions or more. Small ones have from a few hundred million to a few billion: Microsoft’s Phi-3-mini, for example, has 3.8 billion, and its authors described it as small enough to run on a phone. There is no official line between small and large, and the line keeps moving as methods improve.
How a small model gets good#
Three ideas do most of the work. Careful training data: a smaller model trained on cleaner, better-chosen text can match a larger one on some tasks. Distillation: a small model learns to imitate a bigger one. And quantisation: storing the numbers with fewer bits, so the model takes less memory and runs faster. A final step, fine-tuning on an organisation’s own examples, can make a small model very good at one job.
What small models cannot do#
Fewer parameters means less room for facts and for long chains of reasoning. A small model may be fluent and still get a fact wrong, struggle with a complicated problem, or lose track of a long document. They shine at focused, repeated tasks such as classifying, extracting, summarising or answering within a known topic, and are often paired with a larger model for the hard cases.
Where it is used#
- Phones and laptops Writing help, summaries and search that work offline.
- Devices in the field Reading a form, captioning a photo or routing a message on the device itself.
- Private data Working through sensitive text without sending it to a provider.
- High-volume tasks Classifying thousands of messages cheaply.
Why are small models getting attention now?#
Three pressures meet. Running a large model for every request is expensive, and the bill grows with use. Many tasks do not need a model that knows everything. And phones, laptops and cars now carry chips with a neural processing unit that can run a compact model in real time. Together these make a model that is good enough, cheap and local more attractive than the biggest one for many jobs.
There is also a trust reason: a model that runs on the device can work without a connection and does not send private text elsewhere, which matters in hospitals, factories and homes.
Small language models compared with large ones#
A large model knows more and reasons further, but costs more per answer, needs a connection and usually sends your text to a provider. A small model is cheaper and faster and can stay private, but knows less.
The practical question is whether the job is narrow and repeated. If it is, try a small model first and measure it against a large one on your own examples.
| Compared on | Small model | Large model |
|---|---|---|
| Runs on | A phone, a laptop or one local server | Data-centre hardware |
| Cost of each answer | Low | Higher |
| Privacy | Text can stay on the device | Usually sent to a provider |
| Breadth of knowledge | Narrower | Broader |
| Complex reasoning | Weaker, and improving | Stronger |
| Best for | Focused, repeated tasks | Open-ended and hard tasks |
How do you choose between small and large?#
Write down the job and collect a few dozen real examples with the answers you want. Run them through a small model and a large one and compare. If the small model is nearly as good, the saving in cost, delay and privacy usually decides it. If it fails on a few hard cases, send those to a larger model and keep the rest local.
Repeat the test whenever a new small model appears. The best of them improve quickly, and last year’s answer may no longer hold.
Key terms#
- Parameter
- One of the numbers a model adjusts while it learns; a rough measure of its size.
- Distillation
- Training a small model to imitate a larger one.
- Fine-tuning
- Training a model a little further on examples from one task or one organisation.
- On-device model
- A model that runs on the user’s own phone or computer instead of on a remote server.
Common questions#
Can a small model run without the internet?
- Yes. Once the model file is on the device it can run offline, which is a main reason to choose one.
Are small models more private to use?
- Usually. The text can stay on the device, so nothing is sent to a provider. How private it really is depends on the app built around the model.
Will small models replace large ones?
- Not for everything. Small models suit focused jobs; large models still do better on open-ended and hard problems. Many products use both.
Sources and further reading#
Independent pages we checked while writing this guide. They are not Reinnder products, and Reinnder is not affiliated with them.
Reinnder’s angle
How Reinnder looks at small language models
Reinnder Silicon is planned around chips and edge computers that run compact models on the machine itself.
This guide explains the technology in general terms. It is not advice, and it does not describe a Reinnder product on sale.