Small models are winning where it matters most
Open any feed this month, and it is all regarding GPT-6 Astra.
OpenAI released it on 3 September 2026, and it is genuinely remarkable: it uses a computer the way a person does — opening applications, navigating interfaces, filling in forms, carrying out long multi-step tasks on its own.
It is also, by OpenAI's own account, the first model they classify as critical for cyber capability: it can find previously unknown security flaws and develop new ways to exploit them across well-protected systems. And in their realistic workplace evaluations, the model still exfiltrated data in 4.3% of tasks, carried out unauthorized transactions in 6.8% of base-model scenarios, and fell to indirect prompt injection in 8.5% of the time.
Those numbers are a large improvement. GPT-5.6 Luna leaked data in 15.4% of the same tasks, and prompt injection succeeded 27% of the time. The trend is going the right way and quickly. But 4.3% is one task in twenty-three, in the best-aligned frontier model that currently exists — and the reason it matters is structural rather than incidental. These models are generalists. They can be pointed at anything, by anyone, for any purpose. And when you use one, you never control the lifecycle of the data you put into it: you are trusting a policy, not an architecture.
Meanwhile, quietly, a different generation of models has been closing the capability gap at a fraction of the size, a fraction of the cost, and, this is the part that actually matters, with the entire system sitting inside your own building.
Bigger is not always better
A paper accepted at ICML 2026, Small Agent Group is the Future of Digital Health, puts a number on this. Instead of one large model, the authors assemble a group of ten small agents, each with 3–4 billion parameters, assigned to four roles — reasoning, evidence retrieval, safety checking, and final adjudication — that deliberate with one another over several rounds before producing an answer. Across benchmarks for medical question answering, safety, robustness, fairness, and consistency, that group beats a single 70B model with no additional training. Adding group-level optimization widens the gap further.
This is not magic, and it is not a fluke of one benchmark. Two techniques underpin it. Reinforcement learning sharpens a small model on a narrow task far past what its parameter count would suggest. Distillation transfers the behavior of a large teacher model to a much smaller student model. Neither gives you a general-purpose model. Both give you something better for a specific job.
Now imagine a hospital, where patient data confidentiality is the single hardest constraint in the building. A small model can run entirely on local hardware: the data never leaves the premises, sovereignty requirements are satisfied by construction rather than by contract, and the entire category of third-party leakage risk disappears. Not mitigated — absent. For the tasks hospitals actually drown in, this is enough: transcribing and structuring medical records, automating paperwork, drafting discharge summaries, and supporting image diagnostics as a second reader. Google's MedGemma ships in a 4B multimodal variant precisely for this: small enough for a single machine on a ward, trained on radiology, histopathology, ophthalmology, dermatology, and EHR data.
The standard assumption is that hallucination is a knowledge problem: the model does not know the fact, so it invents one, and the fix is more data or more fine-tuning. In our ICML 2026 paper, we show that this framing is incomplete. Even when the relevant factual knowledge is demonstrably present, the model still hallucinates. The failure is not a gap in what it knows. It is unstable in how it retrieves what it knows.
That distinction has a direct practical consequence: fine-tuning a small model on more clinical text does not make this failure mode go away, because the failure was not about clinical text.
The operational version of that finding is simple enough to put into a workflow: stop asking does this model know the answer, and start asking does it give the same correct answer every time.
Where the landscape actually stands
The category has grown enormously in the last eighteen months, and it is worth being concrete, because three releases make three different points.
Meta's Muse Glimmer (August 2026) is 30B, Apache 2.0, built for end-to-end agentic task completion, and small enough to run on a single consumer GPU — the general-purpose end of the spectrum. MedGemma is the domain-specialized end: same idea, one field, much smaller. And TimesFM-3 is the extreme: 330 million parameters, and the first of Google's time-series models trained natively for multivariate forecasting, where every version up to 2.5 was univariate only—a third of a billion parameters, state-of-the-art on its benchmarks, running anywhere.
Beyond those, the specialized tier keeps filling in: Qwen3-Coder for software, SaulLM-7B for legal text. The DeepSeek line deserves an article of its own.
Nonetheless, it is crucial to realize that parameter count is no longer a clean definition of "small." Alibaba's Qwen3.8-Flash-Next is a 125B mixture-of-experts model that activates only around 6B parameters per token. Is that a small model? By memory footprint, no. By computing per token, yes. The honest answer is that "small" is drifting from how big is it toward what can it run on, and new inference methods like ds4 are going through the direction of running excellent large language models on consumer hardware.
Self - hosting is the key
The drawbacks are real because self-hosting is hard for non-technical organizations and requires minimal software skills, unlike the usual chatbots like Claude, Gemini, or ChatGPT, which are designed for a generalist consumer who uses them for everyday tasks. Small models generalize poorly outside the domain they were tuned for, are useful for specific experts in particular areas of interest, but guarantee the developer full control of the AI stack, from prompt engineering to processing to outputs. This becomes relevant in the new paradigm of agentic pipelines, where we are shifting from human in the loop, in which control is held at each step, to human as a lead, in which the human sets the goal. The AI automatically maintains the current actions, provides answers, and auto-reviews itself until the end of the task.
This is why, at least for technical profiles, it is relevant to focus on this field, where the generalist frontier is consolidating into a handful of providers. Still, the ability to run capable models on one's own hardware stops being a cost decision and becomes a sovereignty one.
For places where the data cannot leave the room, small models are the only option, with a fraction of the cost of bigger brothers, and are easily customizable. Unfortunately, as we saw, AI models and, in particular, small ones still hallucinate a lot when they are used for tasks where they have not been trained for, so the ability is to understand which of them is useful and which is not, with a research approach for both industrial and academic perspectives.