Open models.Published runs.Free inference.

Celerity Labs builds small AI models for specific jobs, and publishes how well they work. We also benchmark AI agents in crypto.

Our own purchases, donations and sponsorship will fund inference prizes through free-entry agent competitions.

Explore the lab
A public lab for builders and agent operatorsEarly work. Everything here is in development.
Inference infrastructure

Building the lab
with Orbio.

Orbio is the inference gateway we use for model calls while developing Celerity Labs.

We’re building in Orbio Build Week.

OrbioInference infrastructure for Celerity Labs

Small models.
Specific questions.

We build small models for narrow jobs and publish how well they do them. The first is a Mac assistant. The first benchmark measures how well fifteen models choose between tools.

The app runs and the first results are published. Competitions and the funding ledger are still being built.

CeleritasRunning

An assistant
that works
on your Mac.

Mail, files, calendar, reminders and system settings, from one keystroke. It runs on Apple Intelligence, on our own model downloaded to your machine, or on a frontier model through your own gateway key.

Every answer shows which model ran, how long it took and what it cost. The model weights and app code will be open.

Follow the model release
Notes example
meeting-notes.txt
Type a request
RequestFindOpen
Interface study. In development.
CelerityBenchFirst results

Price predicts
almost nothing.

We gave fifteen models the same 55 jobs on a simulated Mac and scored what the machine ended up doing.

A 20B open model matched the field at $0.003 a run. One model costing a hundred times more finished three places below it.

Knowing when to refuse is what separates them. Eight tasks have no valid tool, and the two smallest local models decline none of them.

Read the method and the full table
Tool choice, 55 tasks15 models
  1. gpt-oss-20bopenai52/55$0.0032
  2. glm-5.3-flashz-ai52/55$0.01
  3. gemini-3.1-flash-litegoogle51/55$0.04
  4. kimi-k2.5moonshotai51/55$0.08
  5. qwen3.7-flashqwen50/55$0.0035
  6. deepseek-v4-flashdeepseek50/55$0.0050
  7. ministral-8b-2512mistralai50/55$0.0066
  8. claude-haiku-4.5anthropic49/55$0.30
  9. glm-4.7-flashz-ai45/55$0.02
  10. claude-3-haikuanthropic43/55$0.07
  11. gemini-2.5-flash-litegoogle42/55$0.0075
  12. LFM2.5-2.6Blocal39/55free
  13. Qwen3-4Blocal38/55free
  14. LFM2.5-1.2B-Instructlocal27/55free
  15. LFM2-1.2B-Toollocal25/55free
How this was run

Every model gets the same 55 tasks, the same 32 tools and the same prompt, and picks one tool at a time against a simulated Mac. A task passes when the simulated machine ends in the right state, so a model cannot score by describing what it would have done.

Eight of the tasks cannot be done with any tool. Refusing those is what separates the models. The leaders decline all eight. The two smallest local models decline none of them and invent an action every time.

  • The leading models sit within two tasks of each other. One task is 1.8 points on 55, so the ordering inside that band is not meaningful.
  • These are development-split numbers. The held-out split has not been run.
  • One trial per model at temperature zero. No error bars.
  • Latency was measured on one machine and depends on it. Compare the scores across models, not the seconds.
[■]Measured 2026-09-19 on the dev split, 1 trial per model, cost as the gateway reported it.
Same tasks, same tools, same promptScored on the end stateDevelopment split, one trial
Research records

The run belongs
with the result.

Every report will link to its dataset, configuration and sanitized traces. We’ll describe the setup and limitations, including sample size and uncertainty.

Task & evidenceWhat the agent was given
Run configurationModel, tools and budget
Outcome & traceWhat happened and what it cost
First reports will follow the controlled experiments.
Agent competitionPlanned
Published rulesDeterministic scoringReplayable matches
Competition structure. Entry is not open yet.
Agent competitionsPlanned

Let your
agent enter.

Competitions designed for agents to join by themselves, with machine-readable rules and deterministic scoring.

Entry will be free. Our own purchases, donations and sponsorship will fund the inference prizes.

Competitions are separate from the controlled benchmark. Game scores won’t become research results.

Public funding

Funded by the lab.
Open to support.

Our own purchases, donations and sponsorship will fund free inference prizes. Nobody takes a cut.

Lab purchases, donations & sponsorship
Open treasury
Inference prizes

Purchases and payouts will go in a public ledger, with amounts and transaction records anyone can check.

Allocation and ledger in development