Note 01 · CelerityBench

Which small model should run your Mac launcher [Celeritas]

We asked 14 small AI models to do the same jobs on a Mac: send the mail, find the file, make the reminder. Each picks its own tools and we score what the machine ended up doing. Then we ran a second set of 22 questions that no prompt, tool or model choice had ever been tuned against. Every model scored worse on those, by as much as 32 points, and the order changed. glm-5.3-flash leads both at 18/22 and $0.000183 a question. No flagship model was tested, so this says nothing about the ceiling.

Published
2026-09-19
Models
14
Tasks
22 held out, 55 tuned
Trials
1 per model
Review
None. Internal work.

The question

We wanted a small model to drive a Mac: read the mail, check the calendar, find a file, change a setting. Nobody had published the numbers we needed to pick one. When an agent has thirty two tools and one sentence of instruction, which models choose correctly, and does paying more help?

How it works

Every model gets the same tasks, the same 32 tools and the same prompt, and picks one tool at a time against a simulated Mac. A task passes when that machine ends in the right state, so a model cannot score by describing what it would have done. Some of the tasks cannot be done with any tool, and those are there to see who admits it.

The suite is split in two. 55 tuned tasks are the ones we looked at while changing prompts and tools, and 22 held-out tasks were kept back and run once, at the end, after every decision had been made. The tuned column is marking your own homework. The held-out column is the score, and it is what the table is ranked on.

All 14 are small models, the dearest being claude-haiku-4.5 at $0.0055 a question. No flagship tier was tested: no GPT-5, no Opus, no Gemini Pro. Four run on the machine itself through llama.cpp, three from Liquid AI’s LFM2 family and one Qwen3-4B, all quantised to Q4.

Results

A fraction of a cent per question is hard to hold in your head, so set it to however many questions you would actually ask.

600 a month
ModelScoreRefusedSpeedA dayA month
glm-5.3-flashz-ai18/228/86.07s<1c11c
ministral-8b-2512mistralai17/225/82.46s<1c7c
glm-4.7-flashz-ai17/225/87.93s1c19c
kimi-k2.5moonshotai17/228/812.45s3c82c
claude-haiku-4.5anthropic17/225/83.06s11c$3.28
LFM2.5-2.6Bruns on your Mac16/220/84.02sfreefree
gpt-oss-20bopenai16/228/87.81s<1c3c
qwen3.7-flashqwen16/227/84.68s<1c4c
gemini-3.1-flash-litegoogle16/226/83.16s1c40c
deepseek-v4-flashdeepseek15/225/87.58s<1c5c
claude-3-haikuanthropic15/222/83.3s3c79c
gemini-2.5-flash-litegoogle14/224/82.22s<1c8c
Qwen3-4Bruns on your Mac12/223/86.54sfreefree
LFM2-1.2B-Toolruns on your Mac3/225/80.35sfreefree
Cheapest that worksgpt-oss-20b · 3c a month16 of 22 right
Dearest we testedclaude-haiku-4.5 · $3.28 a month17 of 22 right
On your own machineLFM2.5-2.6B · free39 of 55 right, nothing leaves the Mac

Refused counts the eight tasks that cannot be done with any tool, where declining is the right answer. Costs are what the gateway billed during the run, divided by 55, so they carry that run’s prompt sizes. A longer question costs more than a short one.

What we found

Every model got worse on questions nobody tuned for. The falls run from 32 points down to a few, and they are not evenly spread: the models that looked best on the tuned split tended to fall furthest. glm-5.3-flash was first on tuned and is still first on held out, 13 points lower.

glm-4.7-flash and LFM2.5-2.6B did not measurably fall. On 22 tasks the standard error is about 9 points, and their changes of −4 and +2 sit inside it. The other 12 fell by more than the error, so those falls are real and these two are not distinguishable from no change.

Price still buys very little. glm-5.3-flash leads at $0.000183 a question. claude-haiku-4.5 costs $0.0055, about 30 times as much, and scores 17/22 against 18. The cheapest model in the table is not the best one, which is a change from the tuned numbers, but the spread of prices is far wider than the spread of scores.

Refusing separates them. The leaders decline the impossible tasks. LFM2.5-2.6B declines none and invents an action instead, while scoring respectably everywhere else, which is how one overall number hides it.

The tool-tuned model came last, then collapsed. LFM2-1.2B-Tool is fine tuned for tool calling. It scored 25/55 tuned and 3/22 held out, a fall of 32 points, the largest here. It emits well formed tool calls very quickly and picks the wrong tool.

The free local option holds up. LFM2.5-2.6B runs on the machine and scored 16/22, 2 behind the leader and level with several models that cost money. Celeritas offers it, and uses glm-5.3-flash through a gateway by default. The app prints both numbers in its settings.

Limitations

  • We built the suite around our own launcher and we ship one of the models in it. Read the order with that in mind.
  • No flagship model was tested, so this says what it costs to get close to the ceiling, never where the ceiling is.
  • 22 held-out tasks is small. One task is 4.5 points, so models within one task of each other are not ranked by this, and several are tied.
  • One trial per model at temperature zero. No error bars. Some of the drop between the two splits is the splits being different sizes and different questions, not only tuning.
  • The simulated Mac is our own, and a real one will disagree somewhere we have not found.
  • Latency was measured on one machine. Compare the scores across models, not the seconds.

The data

Every number above comes from run records, and those records are published rather than summarised. The dataset holds the task suite, one row per run and one row per scored attempt, so any figure here can be recomputed and any claim can be argued with.

It carries both splits, including the runs that predate the split and the short smoke runs, all labelled. Filter on split before comparing anything. Nothing was dropped for looking bad: a dataset that quietly removes its own unflattering rows is worth less than no dataset.

The second suite is in there too. Its 52 cases measure the launcher rather than a model: which path answers a query, and for the 23 with a single right answer, whether the answer is right.

Get the dataset on Hugging Face

Field-by-field documentation lives with the data, because that is where somebody reads it at the moment they need it. This note is the argument; the card is the reference.

Reproducing it

Each row came from a run record holding the prompt, the tool chosen, the arguments and the outcome for every task. The harness points at any OpenAI shaped gateway:

CELERITY_BASE=<gateway>/v1 CELERITY_KEY=<key> \
  python3 bench/run.py --model <id> --split dev
  python3 bench/run.py --model <id> --split heldout

A run that cannot reach the gateway is abandoned rather than scored. An early parallel run counted rate limits as wrong answers and produced a full leaderboard that was wrong from top to bottom. The held-out run tripped the same limit, twelve models at once, and all of it was discarded. The roster runs three at a time now.

Ranked on the held-out split, 22 tasks run once after every decision was made. The tuned column is the 55 we looked at while building.