The question
We wanted a small model to drive a Mac: read the mail, check the calendar, find a file, change a setting. Nobody had published the numbers we needed to pick one. When an agent has thirty two tools and one sentence of instruction, which models choose correctly, and does paying more help?
How it works
Every model gets the same tasks, the same 32 tools and the same prompt, and picks one tool at a time against a simulated Mac. A task passes when that machine ends in the right state, so a model cannot score by describing what it would have done. Some of the tasks cannot be done with any tool, and those are there to see who admits it.
The suite is split in two. 55 tuned tasks are the ones we looked at while changing prompts and tools, and 22 held-out tasks were kept back and run once, at the end, after every decision had been made. The tuned column is marking your own homework. The held-out column is the score, and it is what the table is ranked on.
All 14 are small models, the dearest being claude-haiku-4.5 at $0.0055 a question. No flagship tier was tested: no GPT-5, no Opus, no Gemini Pro. Four run on the machine itself through llama.cpp, three from Liquid AI’s LFM2 family and one Qwen3-4B, all quantised to Q4.
Results
A fraction of a cent per question is hard to hold in your head, so set it to however many questions you would actually ask.
| Model | Score | Refused | Speed | A day | A month | |
|---|---|---|---|---|---|---|
| glm-5.3-flashz-ai | 18/22 | 8/8 | 6.07s | <1c | 11c | |
| ministral-8b-2512mistralai | 17/22 | 5/8 | 2.46s | <1c | 7c | |
| glm-4.7-flashz-ai | 17/22 | 5/8 | 7.93s | 1c | 19c | |
| kimi-k2.5moonshotai | 17/22 | 8/8 | 12.45s | 3c | 82c | |
| claude-haiku-4.5anthropic | 17/22 | 5/8 | 3.06s | 11c | $3.28 | |
| LFM2.5-2.6Bruns on your Mac | 16/22 | 0/8 | 4.02s | free | free | |
| gpt-oss-20bopenai | 16/22 | 8/8 | 7.81s | <1c | 3c | |
| qwen3.7-flashqwen | 16/22 | 7/8 | 4.68s | <1c | 4c | |
| gemini-3.1-flash-litegoogle | 16/22 | 6/8 | 3.16s | 1c | 40c | |
| deepseek-v4-flashdeepseek | 15/22 | 5/8 | 7.58s | <1c | 5c | |
| claude-3-haikuanthropic | 15/22 | 2/8 | 3.3s | 3c | 79c | |
| gemini-2.5-flash-litegoogle | 14/22 | 4/8 | 2.22s | <1c | 8c | |
| Qwen3-4Bruns on your Mac | 12/22 | 3/8 | 6.54s | free | free | |
| LFM2-1.2B-Toolruns on your Mac | 3/22 | 5/8 | 0.35s | free | free |
Refused counts the eight tasks that cannot be done with any tool, where declining is the right answer. Costs are what the gateway billed during the run, divided by 55, so they carry that run’s prompt sizes. A longer question costs more than a short one.
What we found
Every model got worse on questions nobody tuned for. The falls run from 32 points down to a few, and they are not evenly spread: the models that looked best on the tuned split tended to fall furthest. glm-5.3-flash was first on tuned and is still first on held out, 13 points lower.
glm-4.7-flash and LFM2.5-2.6B did not measurably fall. On 22 tasks the standard error is about 9 points, and their changes of −4 and +2 sit inside it. The other 12 fell by more than the error, so those falls are real and these two are not distinguishable from no change.
Price still buys very little. glm-5.3-flash leads at $0.000183 a question. claude-haiku-4.5 costs $0.0055, about 30 times as much, and scores 17/22 against 18. The cheapest model in the table is not the best one, which is a change from the tuned numbers, but the spread of prices is far wider than the spread of scores.
Refusing separates them. The leaders decline the impossible tasks. LFM2.5-2.6B declines none and invents an action instead, while scoring respectably everywhere else, which is how one overall number hides it.
The tool-tuned model came last, then collapsed. LFM2-1.2B-Tool is fine tuned for tool calling. It scored 25/55 tuned and 3/22 held out, a fall of 32 points, the largest here. It emits well formed tool calls very quickly and picks the wrong tool.
The free local option holds up. LFM2.5-2.6B runs on the machine and scored 16/22, 2 behind the leader and level with several models that cost money. Celeritas offers it, and uses glm-5.3-flash through a gateway by default. The app prints both numbers in its settings.
Limitations
- We built the suite around our own launcher and we ship one of the models in it. Read the order with that in mind.
- No flagship model was tested, so this says what it costs to get close to the ceiling, never where the ceiling is.
- 22 held-out tasks is small. One task is 4.5 points, so models within one task of each other are not ranked by this, and several are tied.
- One trial per model at temperature zero. No error bars. Some of the drop between the two splits is the splits being different sizes and different questions, not only tuning.
- The simulated Mac is our own, and a real one will disagree somewhere we have not found.
- Latency was measured on one machine. Compare the scores across models, not the seconds.
The data
Every number above comes from run records, and those records are published rather than summarised. The dataset holds the task suite, one row per run and one row per scored attempt, so any figure here can be recomputed and any claim can be argued with.
It carries both splits, including the runs that predate the split and the short smoke runs, all labelled. Filter on split before comparing anything. Nothing was dropped for looking bad: a dataset that quietly removes its own unflattering rows is worth less than no dataset.
The second suite is in there too. Its 52 cases measure the launcher rather than a model: which path answers a query, and for the 23 with a single right answer, whether the answer is right.
Get the dataset on Hugging Face
Field-by-field documentation lives with the data, because that is where somebody reads it at the moment they need it. This note is the argument; the card is the reference.
Reproducing it
Each row came from a run record holding the prompt, the tool chosen, the arguments and the outcome for every task. The harness points at any OpenAI shaped gateway:
CELERITY_BASE=<gateway>/v1 CELERITY_KEY=<key> \ python3 bench/run.py --model <id> --split dev python3 bench/run.py --model <id> --split heldout
A run that cannot reach the gateway is abandoned rather than scored. An early parallel run counted rate limits as wrong answers and produced a full leaderboard that was wrong from top to bottom. The held-out run tripped the same limit, twelve models at once, and all of it was discarded. The roster runs three at a time now.