Reinnder LabsNow · 2026 to 2027

What are small language models? Compact AI, explained

A small language model is an AI model that understands and writes text but is compact enough to run on a phone, a laptop or one local machine, rather than needing a data centre.

Now Reinnder Silicon Updated 4 October 2026 · 4 min read

Stage
Now · 2026 to 2027
Area of work
Reinnder Silicon
Used in
Phones and laptops · Devices in the field
Key terms
4 defined · glossary

Figure · Illustration

Four ways to make a small model capable

  1. Train on carefully chosen, high-quality text, so that a small model learns more from less.

  2. Teach the small model to imitate the answers of a larger one.

  3. Store the numbers with fewer bits, so the model fits in less memory and runs faster.

  4. Train a little further on examples from one task, so the model is sharp at that job.

Illustration of common techniques. Real models usually combine several of them.

What counts as small#

Size is measured in parameters, the numbers a model learns. The largest models have hundreds of billions or more. Small ones have from a few hundred million to a few billion: Microsoft’s Phi-3-mini, for example, has 3.8 billion, and its authors described it as small enough to run on a phone. There is no official line between small and large, and the line keeps moving as methods improve.

How a small model gets good#

Three ideas do most of the work. Careful training data: a smaller model trained on cleaner, better-chosen text can match a larger one on some tasks. Distillation: a small model learns to imitate a bigger one. And quantisation: storing the numbers with fewer bits, so the model takes less memory and runs faster. A final step, fine-tuning on an organisation’s own examples, can make a small model very good at one job.

What small models cannot do#

Fewer parameters means less room for facts and for long chains of reasoning. A small model may be fluent and still get a fact wrong, struggle with a complicated problem, or lose track of a long document. They shine at focused, repeated tasks such as classifying, extracting, summarising or answering within a known topic, and are often paired with a larger model for the hard cases.

Where it is used#

  • Phones and laptops Writing help, summaries and search that work offline.
  • Devices in the field Reading a form, captioning a photo or routing a message on the device itself.
  • Private data Working through sensitive text without sending it to a provider.
  • High-volume tasks Classifying thousands of messages cheaply.

Why are small models getting attention now?#

Three pressures meet. Running a large model for every request is expensive, and the bill grows with use. Many tasks do not need a model that knows everything. And phones, laptops and cars now carry chips with a neural processing unit that can run a compact model in real time. Together these make a model that is good enough, cheap and local more attractive than the biggest one for many jobs.

There is also a trust reason: a model that runs on the device can work without a connection and does not send private text elsewhere, which matters in hospitals, factories and homes.

Small language models compared with large ones#

A large model knows more and reasons further, but costs more per answer, needs a connection and usually sends your text to a provider. A small model is cheaper and faster and can stay private, but knows less.

The practical question is whether the job is narrow and repeated. If it is, try a small model first and measure it against a large one on your own examples.

Small and large language models compared
Compared onSmall modelLarge model
Runs onA phone, a laptop or one local serverData-centre hardware
Cost of each answerLowHigher
PrivacyText can stay on the deviceUsually sent to a provider
Breadth of knowledgeNarrowerBroader
Complex reasoningWeaker, and improvingStronger
Best forFocused, repeated tasksOpen-ended and hard tasks

How do you choose between small and large?#

Write down the job and collect a few dozen real examples with the answers you want. Run them through a small model and a large one and compare. If the small model is nearly as good, the saving in cost, delay and privacy usually decides it. If it fails on a few hard cases, send those to a larger model and keep the rest local.

Repeat the test whenever a new small model appears. The best of them improve quickly, and last year’s answer may no longer hold.

Key terms#

Parameter
One of the numbers a model adjusts while it learns; a rough measure of its size.
Distillation
Training a small model to imitate a larger one.
Fine-tuning
Training a model a little further on examples from one task or one organisation.
On-device model
A model that runs on the user’s own phone or computer instead of on a remote server.

Common questions#

Can a small model run without the internet?

Yes. Once the model file is on the device it can run offline, which is a main reason to choose one.

Are small models more private to use?

Usually. The text can stay on the device, so nothing is sent to a provider. How private it really is depends on the app built around the model.

Will small models replace large ones?

Not for everything. Small models suit focused jobs; large models still do better on open-ended and hard problems. Many products use both.

Sources and further reading#

Independent pages we checked while writing this guide. They are not Reinnder products, and Reinnder is not affiliated with them.

Reinnder’s angle

How Reinnder looks at small language models

Reinnder Silicon is planned around chips and edge computers that run compact models on the machine itself.

Reinnder Silicon

This guide explains the technology in general terms. It is not advice, and it does not describe a Reinnder product on sale.