Ling-3.0-Tiny: Agentic AI on the worst computer you own
TL;DR
- Ling-3.0-Tiny brings Summer 2025 frontier intelligence to every Apple Silicon Mac on the market
- It’s below larger open-weight models overall but offers exceptional, meticulous web research performance—beating Sonnet 5 in my tests
In 2024, a popular question to stump LLMs was “How many R’s are in the word ‘strawberry’?”
Two years later, that question can now be answered by a model that can run on your worst computer.
Ling-3.0-Tiny is an open-weight model developed by the Ant Group that squeezes formerly frontier intelligence into a shockingly small format. On Artificial Analysis’ Intelligence index, it gets a score of 25, one point away from Google’s Gemma 4 26B, and 5 behind the dense model Gemma 4 31B—Google’s best open-weight releases. For a better point of comparison: the cutting edge of August 2025, ChatGPT-5, earns a 15.
Model performance escalated quickly from last summer’s highs, but the incredible thing here is how little it takes a year later to match that performance. Ling-3.0-Tiny is a 7.9 billion-parameter Mixture-of-Experts model, so only 1.3B loads at a time, making it fast on practically any hardware.
Token Generation Performance
I opted to run the Bartowksi Q6_K quant from Hugging Face. At its full 128k context window, it used under 8 GB of RAM on my 32 GB RAM MacBook Air—a smaller 4-bit release at full context could quite likely be run on a $699 Apple Neo at reasonable speeds.
On my laptop, it handily beats the speeds of the 3x larger Gemma 4 26B:
Note on the table below:
- tg = “token generation.” This gets slower as the context window fills up.
- pp512 = prompt prefill performance with a 512-token prompt.
| Model | cold tg | tg @ 16K context | pp512 | Memory |
|---|---|---|---|---|
| Ling-3.0-Tiny (7.9B), Q4_K_M | 60.4 | 43.8 | 911 | 5.85 GiB |
| Ling-3.0-Tiny (7.9B), Q6_K | 53.8 | 44.1 | 880 | 7.64 GiB |
| Gemma 4 26B-A4B QAT, MLX 0.31.3 | 36.6 | 25.2 | 317 | 17.12 GiB |
| Gemma 4 26B-A4B QAT, llama.cpp | 36.2 | 30 | 354 | 18.8 GiB |
Ling-3.0-Tiny’s Greenwald Eval
On my private test suite, run with the Pi harness, Ling-3.0-Tiny passed just 4 cases: 3 code writing and one web research. This makes it the weakest model I’ve tested, though that’s a bit misleading, as we’ll discuss below.
Despite its superior token-generation, it lost to Gemma 4 26B on wall-clock—it needed to think harder to get to the same result on 3 of 4 tests, using many more tokens and ultimately taking longer.
| Case | Ling toks (in/out) | Ling wall-clock | Gemma toks (in/out) | Gemma wall-clock |
|---|---|---|---|---|
| code-writing/flashcard-parser-lite | 19,853 / 16,172 | 460s | 4,086 / 3,542 | 127s |
| sql-authoring/habit-completion-window | 2,603 / 1,535 | 38s | 1,589 / 4,709 | 151s |
| swift-authoring/streak-count | 5,012 / 4,633 | 112s | 1,332 / 2,585 | 85s |
| web-research/llamacpp-flag-lookup | 30,079 / 4,563 | 256s | 10,415 / 1,074 | 78s |
Gemma 4 26B is also a little smarter in some areas: it passed two more of my 16-and-counting evals, including a more complex coding task and an objectivity test to stand its ground on pushback. That 1-point difference is a little wider than the index benchmarks suggest.
But the big win here is that a model of this size can do any of this work at all, and the RAM footprint means it can run always-on without closing my browser or handing over the whole computer to LLM performance like I need to do with models in the 25B+ parameter category.
Like other models of recent weeks, it follows the pattern of thinking longer to make up for lack of raw intelligence. That thinking capability is in fact essential: it failed all of my evals with reasoning turned off.
Real-World Usage
Evals aren’t the only story. Ling-3.0-Tiny is exceptionally good at research—I would say even frontier level.
Given the request “What are the 3 most recent releases by Walrus pedals?,” Ling immediately did web searches and returned the correct list of pedals, release dates, and prices.
By contrast:
- Sonnet 5 at low effort returned the right pedal list and approximate release dates with no prices
- Haiku 4.5 returned 2 of 3 and got the third incorrect
- Gemma 4 26B searched for pedals released in 2025 and returned the wrong list
Gemma 4 26B tends to need more guidance, and was able to return the correct list with a different prompt, after thinking about it for several minutes: “What are the 3 most recent pedals released by Walrus Audio? It is August 25, 2026. Start by looking at the official Walrus website.”
Asked, “Who is music journalist David Greenwald’s favorite band or artist?” it managed to track down the week I wrote about the Softies (indeed, my favorite band) for the site One Week, One Band. It came back with an exact quote and link—once again, meticulously complete.
Haiku 4.5 literally refused the question (“I don’t have information about music journalist David Greenwald’s favorite band or artist.”) and Sonnet 5 at high effort—with and without my “web explorer” skill, which has been my gold standard for research—also didn’t make it this far.
Ling also gave me a great answer about a MySQL index technical question, which it jumped right to web search to answer instead of going from memory.
After about a half-hour of this, I was floored. What else can it do?
The Ant Group’s X account shared some helpful demos, showcasing it for Obsidian text interfacing, translation, and computer use. I have yet to try these, but will report back.
Try Ling-3.0-Tiny via Hugging Face (or grab the bartowski release I used.)