36 AI models against 20 real small-business tasks
Invoices, quotes, rotas, review replies and inbox triage, each with a trap a careless answer falls into. Every answer passes or fails as a whole, and every pass is priced.
A pilot: 20 of the edition’s 40 tasks, 1 run a model.
- Models
- 36
- Tasks
- 20 of 40
- Runs
- 720
- Rubric grader
- Calibrated
Run in October 2026. Data under CC BY 4.0.
The AIGROW small-business benchmark gives 36 AI models 20 everyday office tasks and marks every answer pass or fail against written checks. In the October 2026 edition Qwen3.8 Flash, Kimi K3 and Gemini 3.1 Pro (preview) tied on top at 95%, and gpt-oss-120b gave the cheapest passes at $0.422 per thousand. Every task, check and result is published.
Findings
What the October 2026 edition shows, and how sure it is.
Qwen3.8 Flash, Kimi K3 and Gemini 3.1 Pro (preview) each passed 95% of their runs, the highest rate in the edition.
The 95% intervals of 6 more models reach that rate. At 20 runs a model, they cannot be told apart from the top, and the bars in the leaderboard show how far each could move.
A thousand passes cost from $0.422 to $33.11, a 78× spread.
gpt-oss-120b gave the cheapest passes and Claude Fable 5.1 the dearest. A model that fails often pays for its failures here, because every run it made is in the cost.
Reply publicly to a negative restaurant review without revealing private details was the hardest task: 22% of runs passed.
A task in English from the review replies category. Its page lists the traps and the checks most models failed.
Tasks in Italian passed 83% of runs, against 65% in English.
The Italian tasks are their own tasks, written for Italian businesses with Italian tax and labelling rules, not translations of the English ones. The gap mixes language with difficulty.
Median answers took from 1.2 s on Gemini 3.5 Flash-Lite to 43 s on Qwen3.8 Max.
Measured from the request to the last word of the answer, with failed calls left out. The p90 column shows the slow tail a customer would wait through.
The leaderboard
Every model on every task, sorted by how often its answer passed. Sort by cost to see what a pass costs, or by speed to see who answers first.
Highest success rate first.
| # | Model | Success rate | Per 1,000 passes | Median | p90 | English | Italian |
|---|---|---|---|---|---|---|---|
| 1 | Qwen3.8 FlashAlibaba (Qwen) | 95% interval 76 to 99% | $1.09 | 29 s | 44 s | 92% | 100% |
| 1 | Kimi K3Moonshot AI · open weights | 95% interval 76 to 99% | $18.44 | 15 s | 78 s | 100% | 86% |
| 1 | Gemini 3.1 Pro (preview)Google | 95% interval 76 to 99% | $29.67 | 17 s | 21 s | 92% | 100% |
| 4 | GLM-5.3-FlashZ.ai · open weights | 95% interval 70 to 97% | $0.831 | 15 s | 31 s | 92% | 86% |
| 4 | DeepSeek V4.1 FlashDeepSeek · open weights | 95% interval 70 to 97% | $1.05 | 14 s | 127 s | 85% | 100% |
| 4 | DeepSeek V4 ProDeepSeek · open weights | 95% interval 70 to 97% | $5.14 | 21 s | 49 s | 85% | 100% |
| 4 | GLM-5.3Z.ai · open weights | 95% interval 70 to 97% | $7.02 | 9.2 s | 31 s | 92% | 86% |
| 4 | Gemini 3.8 FlashGoogle | 95% interval 70 to 97% | $7.17 | 11 s | 16 s | 85% | 100% |
| 4 | Grok 4.7xAI | 95% interval 70 to 97% | $9.41 | 17 s | 27 s | 92% | 86% |
| 10 | Qwen3.7 PlusAlibaba (Qwen) | 95% interval 64 to 95% | $3.58 | 29 s | 41 s | 77% | 100% |
| 10 | Claude Sonnet 5Anthropic | 95% interval 64 to 95% | $6.70* | 25 s | 51 s | 77% | 100% |
| 10 | GPT-5.6 SolOpenAI | 95% interval 64 to 95% | $11.33* | 27 s | 37 s | 77% | 100% |
| 10 | Qwen3.8 MaxAlibaba (Qwen) | 95% interval 64 to 95% | $16.32 | 43 s | 82 s | 77% | 100% |
| 10 | Claude Opus 5Anthropic | 95% interval 64 to 95% | $16.89* | 27 s | 83 s | 77% | 100% |
| 10 | GPT-5.5OpenAI | 95% interval 64 to 95% | $17.02* | 22 s | 34 s | 77% | 100% |
| 10 | GPT-6 AstraOpenAI | 95% interval 64 to 95% | $27.91* | 11 s | 38 s | 77% | 100% |
| 10 | Claude Fable 5.1Anthropic | 95% interval 64 to 95% | $33.11* | 26 s | 51 s | 77% | 100% |
| 18 | GPT-5.6 LunaOpenAI | 95% interval 58 to 92% | $0.677* | 28 s | 34 s | 69% | 100% |
| 18 | Muse Glimmer 30BMeta · open weights | 95% interval 58 to 92% | $3.10 | 13 s | 27 s | 69% | 100% |
| 18 | Grok 4.3xAI | 95% interval 58 to 92% | $4.10 | 7.4 s | 9.7 s | 77% | 86% |
| 18 | GPT-5.6 TerraOpenAI | 95% interval 58 to 92% | $6.68* | 29 s | 36 s | 69% | 100% |
| 18 | Qwen3.8 27BAlibaba (Qwen) · open weights | 95% interval 58 to 92% | $7.30 | 43 s | 89 s | 77% | 86% |
| 23 | gpt-oss-120bOpenAI · open weights | 95% interval 53 to 89% | $0.422 | 33 s | 69 s | 77% | 71% |
| 23 | MiniMax M3MiniMax · open weights | 95% interval 53 to 89% | $1.82 | 6.4 s | 45 s | 77% | 71% |
| 23 | Claude Sonnet 4.6Anthropic | 95% interval 53 to 89% | $11.24* | 28 s | 54 s | 62% | 100% |
| 23 | Kimi K2.6Moonshot AI · open weights | 95% interval 53 to 89% | $11.68 | 28 s | 73 s | 69% | 86% |
| 27 | Gemma 4 31BGoogle · open weights | 95% interval 39 to 78% | $0.558 | 11 s | 36 s | 46% | 86% |
| 28 | Llama 4 MaverickMeta · open weights | 95% interval 26 to 66% | $0.665 | 14 s | 21 s | 38% | 57% |
| 28 | Gemini 3.1 Flash-LiteGoogle | 95% interval 26 to 66% | $1.41 | 1.7 s | 3.2 s | 38% | 57% |
| 30 | GPT-5.4 miniOpenAI | 95% interval 22 to 61% | $3.35 | 1.7 s | 2.3 s | 23% | 71% |
| 30 | Mistral Medium 3.5Mistral · open weights | 95% interval 22 to 61% | $7.40 | 1.6 s | 2.3 s | 38% | 43% |
| 32 | Gemini 3.5 Flash-LiteGoogle | 95% interval 18 to 57% | $2.47 | 1.2 s | 1.5 s | 23% | 57% |
| 32 | Claude Haiku 4.5Anthropic | 95% interval 18 to 57% | $6.62 | 2.5 s | 3.4 s | 23% | 57% |
| 34 | GPT-5.4 nanoOpenAI | 95% interval 15 to 52% | $1.51 | 2.7 s | 3.8 s | 23% | 43% |
| 35 | Ministral 3 14BMistral · open weights | 95% interval 8 to 42% | $1.07 | 3.9 s | 5.4 s | 8% | 43% |
| 35 | Mistral Small 4Mistral · open weights | 95% interval 8 to 42% | $1.33 | 2.0 s | 20 s | 15% | 29% |
Success rate is the share of runs that passed, with its 95% interval drawn under it: a model’s true rate could sit anywhere on the bar. Cost per pass divides the cost of every run, failures included, by the runs that passed. An asterisk marks a cost worked from the maker’s list price, where the provider reported none.
By kind of work
The share of runs that passed in each kind of task, across every model, and the models that passed it most often.
| Kind of task | Tasks | Runs passed | Passed most often |
|---|---|---|---|
| Request routing | 2 | 96% | 33 models at 100% |
| Invoice extraction | 2 | 90% | 29 models at 100% |
| Email triage | 1 | 83% | 30 models at 100% |
| Meeting summaries | 1 | 83% | 30 models at 100% |
| Business calculations | 2 | 76% | 19 models at 100% |
| Lead qualification | 1 | 72% | 26 models at 100% |
| Bookkeeping | 2 | 71% | 24 models at 100% |
| Appointment scheduling | 2 | 69% | 21 models at 100% |
| Policy questions | 1 | 69% | 25 models at 100% |
| Customer support | 2 | 67% | 17 models at 100% |
| Product descriptions | 1 | 58% | 21 models at 100% |
| Quotes | 2 | 51% | 9 models at 100% |
| Review replies | 1 | 22% | 8 models at 100% |
The tasks
Hardest first. Each page shows the task as the models saw it, the traps in it, every check in words and what each model got wrong.
How it is marked
The same instructions and material for every model, one answer a run, and a verdict that does not depend on how the answer sounds.
- Pass or fail
- An answer passes only if every fixed check passes. On the tasks that also carry a rubric, it has to meet most of its criteria too, and the criteria that matter most are mandatory. There is no partial credit.
- Fixed checks
- Structured answers are read field by field against the right values, with a tolerance on money and hours. Written answers are checked for what they must say, what they must not say and how long they may be. The same code marks every model.
- The rubric grader
- Rubric criteria are marked by gpt-5-6-sol at temperature zero, each with a reason. Every answer is marked twice, and a third mark settles any criterion the two disagree on. Its own verdict is ignored: code applies the pass rule to the criteria it marked. Before the run it was tested twice on answers whose verdict is known.
- One protocol
- Each run is one instruction message and one message with the material: no tools, no web search, no JSON mode, each model’s default sampling and up to 4,000 tokens of output, reasoning included. An answer cut off at that limit is marked as it stands. 1 run a model on each task.
- Audited rules
- Before an edition is published, every failure on a deciding rule is read. A rule that fails a correct answer for leaving out a detail the instructions never asked for is corrected, and those answers are marked again. A rule that catches a real mistake stays, however many models it fails.
- Cost and time
- A run costs what the provider billed for it, or, where the provider reports no cost, its tokens at the maker’s list price. Time runs from the request to the last word of the answer; calls that failed are left out of it.
- Intervals
- Every success rate carries a Wilson 95% interval. With a few dozen runs a model, a gap of several points between two models can be chance, and the bars say so.
- The data
- Tasks, checks and results are published under CC BY 4.0: cite AIGROW and link this page. Every task page carries the edition’s canary string, so the tasks can be kept out of training sets.
Put the models that pass to work
An AIGROW agent does this kind of work inside your own tools, against rules you approve in writing. To run the models yourself, API credit costs a quarter of list price.
October 2026 edition, pilot