
Open-source LLMs have exploded this year. It’s crazy how far we’ve come in the last few months. Chinese models have matured and can do remarkable things at a fraction of the cost. Even European models have a bit of a renaissance!
But as much as we love our little Qwens and Kimis, they are no match for the big boys. OpenAI and Anthropic still run the really large language models.
It’s a common pattern these days to use a big frontier model for planning and a smaller (local) model for execution. Fable designs the app; Qwen writes the code.
That is already feasible today, and a lot of teams use such an approach to keep the token cost down.
In such a setup, the majority of the work gets done by non-frontier models.
It spells bad news for the big boys. If cheap, fast, interchangeable models can reach Sonnet or Opus levels, companies will do the math and diversify. If only their largest models are used, their market share shrinks. That is starting to happen already.
Today, GPT6 Astra was released. It’s OpenAI’s answer to Fable.
It is rumoured to be a lot better than its predecessors. It breaks all the benchmarks. There is hype and even a commercial.
But benchmarks don’t tell us about its real-world applications. We need to ask ourselves the obvious but painful question:
What can Astra do today that GPT5.6 Sol can’t?
Which use cases are possible today that were impossible just a day ago?
It doesn’t improve my day-to-day coding workflow. It isn’t affordable as a Hermes model. What exactly would I use Astra for?
The answer paints a picture of a plateau where each new frontier model is just incrementally better. Of better numbers but similar value. Of a curve that is flattening. Of diminishing returns to the ever-bigger models.
As modest-sized models catch up with the frontier and become faster and more memory- and energy-efficient, we might be entering a period where they can replace Fable in day-to-day work.
They won’t be as capable, and they won’t break the benchmark.
But they will be good enough and an order of magnitude cheaper.
If we have those, why would we ever use a more expensive frontier model?
