The future of AI is now 3 frontiers

What is going to happen next in AI? I cannot predict the specifics—or the economics, which is out of my pay grade, sorry, investors—but the future transformed this summer. What was a single path forward is now three, and the next few months may be as eventful as the season we’re closing.

Here’s what I think will happen next.

Where we are: the stale frontier

Anthropic’s Fable has been out (and then not out, and then out again) since June. Opus 5 looks better on benchmarks but less so in my practical work, so let’s leave that out for now. OpenAI’s 5.6 models arrived in early July.

The summer since then has been hard to believe. Numerous Chinese labs have released free, open-weight categories across every category:

  • ling-3.0-tiny: A small model that can run on almost any hardware that’s as good as the summer 2025 frontier (see my Ling review here)
  • qwen 3.8 27B: A model benchmarking at Opus levels that is small enough to run—albeit at inchworm pace—on consumer hardware. (This one did not do great in my evals, so I have some doubts here, but I haven’t tried it at the higher effort levels where it is supposed to shine.)
  • DeepSeek v4 Flash 0731: A strong mid-tier model that can run for cents per task in the cloud. DeepSeek, the company, raised its cloud hosting prices, but this is still available at too-cheap-to-meter rates from providers such as DeepInfra.
  • GLM-5.3-Flash: A genuinely Opus-level model which, at its current discount pricing, is cheaper than DeepSeek 0731. During its “stealth” stint as Ox Alpha, I did a week of real work with it: it’s excellent, no ifs or buts.

These are the just the models I’ve tried personally. On the larger and spendier side, there’s also Kimi K3, which allegedly matches Fable; Qwen 3.8 Max, the big brother of the 27B release; and Hy4 preview, which isn’t even its final form. Developer Tencent noted in today’s (!) announcement: “There is real headroom left in both pre-training and post-training, and we are shipping with known issues — among them, spending longer than necessary reasoning through complex tasks, and a tendency to over-verify its own work.”

U.S. companies have also joined the race, with a return from Meta (the DeepSeek-caliber Muse Spark 1.2), a pair of interesting models from Poolside, and X.ai’s Grok.

What’s the AI we actually need?

Calvin French-Owen touched on this in his post Small Models Have Arrived:

Think of the people you interact with on a daily basis: coworkers, vendors, and customers. Nine times out of ten, you want someone who is super responsive, and just handles things for you. Most of the “human tokens” at companies today are spent this way — hiring skews heavily toward the fast/cheap/good-enough archetype.

At this point, let me share some context and some strong opinions as a software engineering manager:

The November 2025 release of Opus 4.5 marked the turning point for many developers on “AI can be helpful but isn’t serious” to “AI can do real work but needs help.” The releases since have built on this, improving the ability for AI to work autonomously and brought us the idea of “long-horizon” tasks.

Opus 4.5 and 4.6 addressed the intelligence level I needed for most of my work but not the autonomy—Opus 4.8 can handle multiple steps, sub-agents, and work over long periods with few mistakes and great attention to detail. It’s the difference between prompt engineering and a software factory.

In my personal time, I am now building working software—for iOS, MacOS, an Obsidian TypeScript plugin—that I do not understand. I am not reading the code. I’m just the engineering manager! I tell Opus 4.8 what to do, it runs through a very extensive workflow process that we have tuned over months, and then 1-2 hours later, most of the time, I get a working feature pull request on the other end.

I don’t need any AI to get any better than this. If you aren’t doing medical research or innovative science, you probably don’t either.

Now, at the start of the summer, June 2026, Opus 4.8 was the best model on earth by a real margin. My only option for serious agentic work was Anthropic or OpenAI.

That moment is over.

The immediate future of AI is a triple frontier

You’ve heard the saying “fast, cheap, good—pick two.” Well. Not anymore.

1. Intelligence

Both Anthropic and OpenAI are said to be preparing their next-generation models. Google is overdue for a serious competitor to Claude and ChatGPT. The American frontier has gone stale enough that the Chinese labs have caught up, and so their next-generation models—the Kimi 3.1s and Qwen 4s—will push ahead, too.

Where we go from here also gets into the domain of science fiction. RSI? AGI? I won’t make any bets here.

This is the frontier we’ve seen playing out, and the one that’s the least interesting for to me, and maybe to you.

2. Speed

We’ve heard “code is no longer the bottleneck” from a lot of people who haven’t watched Claude take 2 hours to finish building and testing a PR. That’s a bottleneck!

The best we’ve had from the frontier models is Opus’ “fast mode”, which roughly doubles the price and cost. It’s a nice option to have, since that’s double the already painful full rate of API pricing, it’s a pass for me.

This is where the biggest change is coming.

There are multiple tracks forward on speed: model size, inference engine innovations, and hardware.

Opus 4.8 quality in May needed a huge model. In August, looking at GLM-5.3-Flash, it does not. The smaller the models get in size, the less memory bandwidth they need, the faster they will run on the hardware that exists today. Meta’s Muse Spark 2.1, which cruises at over 200 tokens/second, is a good example of this in action.

At the same time, every week there is a new llama.cpp pull request or GitHub repo sharing someone’s mad-scientist scheme to squeeze open-weight models onto consumer hardware. The innovation here will no doubt continue.

But the best mover of needles is purpose-built hardware: the technical innovations of Cerebras and Taalas, aiming at thousands of tokens per second, not dozens or hundreds.

OpenAI has partnered with Cerebras to run their best model, 5.6 Sol, at “ultrafast” speeds. Same work, just delivered in more tokens per second. At 750 toks/second, that’s something like 10x the speed of Opus 4.8 at lower effort levels. This is no doubt going to be extremely expensive when it goes into wide release, but it’s the first two thirds of the good, fast, cheap triangle.

The last corner is also on the way. Cerebras makes cheaper models available already: it currently hosts smaller, less intelligent models, such as Gemma 4 31B, for reasonable prices. The other day, I had Gemma summarize 400 files in 8 minutes, a task that was going to take Sonnet all afternoon, and it cost me a couple of bucks.

In September, they’ll be retiring Gemma and putting Qwen 3.8 27B in its place. The same small but mighty Qwen 3.8 27B which rivals Opus’ benchmarks.

What this means in a week from now, we may have an unofficial Opus Ultrafast mode for a dollar a task. If this particular Qwen can’t do it, it won’t be long until there’s a model that can.

That opens up everything. Imagine building a production-ready, feature-complete app in a couple of days. You can also imagine building infinite slop, but that’s a topic for another day.

3. Cost

Anthropic has yet to respond to the cost pressures of the summer, as Kimi, DeepSeek and others chew into their market share, but that seems unlikely to last.

OpenAI has slashed costs on Luna (its cheapest model) and discounted Sol; fears that subscription pricing and particularly to team plan tiers would go away as token use climbed in the spring seem to have eased.

If Anthropic does pull its $100 and $200 subscriptions—still, to be sure, the best deal in town—nothing is stopping many people from flipping over to GLM-5.3-Flash or DeepSeek V4 Flash 0731 on-demand for the same pricing.

DeepSeek has raised its hosting prices from pennies recently, but the model is freely available and anyone can run it, so it doesn’t matter—DeepInfra and other vendors are still running it at bargain-basement pricing. Annual hardware refresh cycles will only improve the price of these older models going forward, even without specialized technologies, and they’ll get faster, too.

Again, just a few months ago, this wasn’t an option. Now it is. If you haven’t tried these models yet, you should.

Of the three frontiers here, I do think this is the one where I think we may see the least movement in the coming months. Most of the new models are running at dollars per task, not pennies, hovering between Claude’s full-price API rates and the subsidy-rates of subscription costs. User demand may crowd out limited capacity and force costs to go up. Someone has to make money on this stuff.

The real step-change here will be when Opus-class open-weight models can run on $2,000 computers and not $10,000 ones. Perhaps we will hit a limit on scaling down GLM-5.3-Flash’s 300 billion parameters into 30 billion—but maybe not. Imagine how good the Ling-3.0-Tiny equivalent of summer 2027 will be: a 100 token-per-second model you can run for the price of your electric bill, if you’re not running it from Cerebras at 3,000 tokens instead.

Closing thoughts

Unlimited, warp speed Opus 4.8. That’s all I want. And that’s what I think we’re going to get, sooner than later.

What am I doing about it now?

  • Writing evals so I can see the limits and quirks of these new models as they land
  • Iterating on my software factory so warp-speed sub-agent runs are productive and not a waste of tokens
  • Trying to be patient

See you in the future.