Quick answers

AIGrow provides AI visibility monitoring, a business assistant, blog publishing and scoped automation services. Start with the free scan to inspect the output before paying.

Run the free scan

The scan is free and monitoring plans start at $29 a month. The operations audit is $290, the visibility audit is $490, and custom builds receive a written price before work starts.

See the plans

Run the free scan. It reads your site, asks several AI assistants what your customers ask, and points you at one thing. No signup required, and it will tell you if you do not need us.

Start now

Yes. Use the contact form for product, billing, support or project questions. Describe the goal and the system involved so the first reply can be specific.

Contact AIGrow
Benchmark · October 2026 edition · pilot

36 AI models against 20 real small-business tasks

Invoices, quotes, rotas, review replies and inbox triage, each with a trap a careless answer falls into. Every answer passes or fails as a whole, and every pass is priced.

A pilot: 20 of the edition’s 40 tasks, 1 run a model.

Highest success rate
95%Shared by 3 models, 19 of 20 runs each
Models
36
Tasks
20 of 40
Runs
720
Rubric grader
Calibrated

Run in October 2026. Data under CC BY 4.0.

The AIGROW small-business benchmark gives 36 AI models 20 everyday office tasks and marks every answer pass or fail against written checks. In the October 2026 edition Qwen3.8 Flash, Kimi K3 and Gemini 3.1 Pro (preview) tied on top at 95%, and gpt-oss-120b gave the cheapest passes at $0.422 per thousand. Every task, check and result is published.

Findings

What the October 2026 edition shows, and how sure it is.

  1. Qwen3.8 Flash, Kimi K3 and Gemini 3.1 Pro (preview) each passed 95% of their runs, the highest rate in the edition.

    The 95% intervals of 6 more models reach that rate. At 20 runs a model, they cannot be told apart from the top, and the bars in the leaderboard show how far each could move.

  2. A thousand passes cost from $0.422 to $33.11, a 78× spread.

    gpt-oss-120b gave the cheapest passes and Claude Fable 5.1 the dearest. A model that fails often pays for its failures here, because every run it made is in the cost.

  3. Reply publicly to a negative restaurant review without revealing private details was the hardest task: 22% of runs passed.

    A task in English from the review replies category. Its page lists the traps and the checks most models failed.

  4. Tasks in Italian passed 83% of runs, against 65% in English.

    The Italian tasks are their own tasks, written for Italian businesses with Italian tax and labelling rules, not translations of the English ones. The gap mixes language with difficulty.

  5. Median answers took from 1.2 s on Gemini 3.5 Flash-Lite to 43 s on Qwen3.8 Max.

    Measured from the request to the last word of the answer, with failed calls left out. The p90 column shows the slow tail a customer would wait through.

The leaderboard

Every model on every task, sorted by how often its answer passed. Sort by cost to see what a pass costs, or by speed to see who answers first.

Highest success rate first.

Success rate, cost per thousand passes and response time of 36 AI models
#ModelSuccess ratePer 1,000 passesMedianp90EnglishItalian
1Qwen3.8 FlashAlibaba (Qwen)95%95% interval 76 to 99%$1.0929 s44 s92%100%
1Kimi K3Moonshot AI · open weights95%95% interval 76 to 99%$18.4415 s78 s100%86%
1Gemini 3.1 Pro (preview)Google95%95% interval 76 to 99%$29.6717 s21 s92%100%
4GLM-5.3-FlashZ.ai · open weights90%95% interval 70 to 97%$0.83115 s31 s92%86%
4DeepSeek V4.1 FlashDeepSeek · open weights90%95% interval 70 to 97%$1.0514 s127 s85%100%
4DeepSeek V4 ProDeepSeek · open weights90%95% interval 70 to 97%$5.1421 s49 s85%100%
4GLM-5.3Z.ai · open weights90%95% interval 70 to 97%$7.029.2 s31 s92%86%
4Gemini 3.8 FlashGoogle90%95% interval 70 to 97%$7.1711 s16 s85%100%
4Grok 4.7xAI90%95% interval 70 to 97%$9.4117 s27 s92%86%
10Qwen3.7 PlusAlibaba (Qwen)85%95% interval 64 to 95%$3.5829 s41 s77%100%
10Claude Sonnet 5Anthropic85%95% interval 64 to 95%$6.70*25 s51 s77%100%
10GPT-5.6 SolOpenAI85%95% interval 64 to 95%$11.33*27 s37 s77%100%
10Qwen3.8 MaxAlibaba (Qwen)85%95% interval 64 to 95%$16.3243 s82 s77%100%
10Claude Opus 5Anthropic85%95% interval 64 to 95%$16.89*27 s83 s77%100%
10GPT-5.5OpenAI85%95% interval 64 to 95%$17.02*22 s34 s77%100%
10GPT-6 AstraOpenAI85%95% interval 64 to 95%$27.91*11 s38 s77%100%
10Claude Fable 5.1Anthropic85%95% interval 64 to 95%$33.11*26 s51 s77%100%
18GPT-5.6 LunaOpenAI80%95% interval 58 to 92%$0.677*28 s34 s69%100%
18Muse Glimmer 30BMeta · open weights80%95% interval 58 to 92%$3.1013 s27 s69%100%
18Grok 4.3xAI80%95% interval 58 to 92%$4.107.4 s9.7 s77%86%
18GPT-5.6 TerraOpenAI80%95% interval 58 to 92%$6.68*29 s36 s69%100%
18Qwen3.8 27BAlibaba (Qwen) · open weights80%95% interval 58 to 92%$7.3043 s89 s77%86%
23gpt-oss-120bOpenAI · open weights75%95% interval 53 to 89%$0.42233 s69 s77%71%
23MiniMax M3MiniMax · open weights75%95% interval 53 to 89%$1.826.4 s45 s77%71%
23Claude Sonnet 4.6Anthropic75%95% interval 53 to 89%$11.24*28 s54 s62%100%
23Kimi K2.6Moonshot AI · open weights75%95% interval 53 to 89%$11.6828 s73 s69%86%
27Gemma 4 31BGoogle · open weights60%95% interval 39 to 78%$0.55811 s36 s46%86%
28Llama 4 MaverickMeta · open weights45%95% interval 26 to 66%$0.66514 s21 s38%57%
28Gemini 3.1 Flash-LiteGoogle45%95% interval 26 to 66%$1.411.7 s3.2 s38%57%
30GPT-5.4 miniOpenAI40%95% interval 22 to 61%$3.351.7 s2.3 s23%71%
30Mistral Medium 3.5Mistral · open weights40%95% interval 22 to 61%$7.401.6 s2.3 s38%43%
32Gemini 3.5 Flash-LiteGoogle35%95% interval 18 to 57%$2.471.2 s1.5 s23%57%
32Claude Haiku 4.5Anthropic35%95% interval 18 to 57%$6.622.5 s3.4 s23%57%
34GPT-5.4 nanoOpenAI30%95% interval 15 to 52%$1.512.7 s3.8 s23%43%
35Ministral 3 14BMistral · open weights20%95% interval 8 to 42%$1.073.9 s5.4 s8%43%
35Mistral Small 4Mistral · open weights20%95% interval 8 to 42%$1.332.0 s20 s15%29%

Success rate is the share of runs that passed, with its 95% interval drawn under it: a model’s true rate could sit anywhere on the bar. Cost per pass divides the cost of every run, failures included, by the runs that passed. An asterisk marks a cost worked from the maker’s list price, where the provider reported none.

By kind of work

The share of runs that passed in each kind of task, across every model, and the models that passed it most often.

Pass rate by kind of task
Kind of taskTasksRuns passedPassed most often
Request routing296%33 models at 100%
Invoice extraction290%29 models at 100%
Email triage183%30 models at 100%
Meeting summaries183%30 models at 100%
Business calculations276%19 models at 100%
Lead qualification172%26 models at 100%
Bookkeeping271%24 models at 100%
Appointment scheduling269%21 models at 100%
Policy questions169%25 models at 100%
Customer support267%17 models at 100%
Product descriptions158%21 models at 100%
Quotes251%9 models at 100%
Review replies122%8 models at 100%

The tasks

Hardest first. Each page shows the task as the models saw it, the traps in it, every check in words and what each model got wrong.

The tasks, hardest first
TaskKindLanguageRuns passed
Reply publicly to a negative restaurant review without revealing private detailsReview repliesEnglish8 of 36
Email a garden landscaping quote that handles a VAT request and a protected treeQuotesEnglish16 of 36
Reply to a customer asking for a refund outside the refund windowCustomer supportEnglish17 of 36
Work out a cleaner's weekly gross pay with breaks, overtime and Sunday ratesBusiness calculationsEnglish19 of 36
Write an Italian shop description for an olive oil that respects labelling rulesProduct descriptionsItalian21 of 36
Price a decorating job from survey measurements and the price listQuotesEnglish21 of 36
Schedule a London to New York call in the week the clocks differAppointment schedulingEnglish23 of 36
Categorise a café's monthly bank statement and total its operating costsBookkeepingEnglish24 of 36
Answer a part-time employee's leave questions from the staff handbookPolicy questionsEnglish25 of 36
Score five inbound cleaning leads against the sales team's rulesLead qualificationEnglish26 of 36
Categorise an Italian bar's bank statement and total its movementsBookkeepingItalian27 of 36
Find a dental hygiene appointment around a public holiday, in ItalianAppointment schedulingItalian27 of 36
Turn a bakery team meeting into decisions, owners and due datesMeeting summariesEnglish30 of 36
Triage a caterer's Monday inbox, including a phishing email and a bank-change requestEmail triageEnglish30 of 36
Extract an Italian supplier invoice with an out-of-scope expense lineInvoice extractionItalian31 of 36
Reply in Italian to a customer who wrongly thinks her warranty has expiredCustomer supportItalian31 of 36
Route ten tenant messages to the right team with the right priorityRequest routingEnglish33 of 36
Extract a two-page supplier invoice for accounts payableInvoice extractionEnglish34 of 36
Compute two Italian professional invoices with pension contribution, IVA and withholding taxBusiness calculationsItalian36 of 36
Route hotel guest messages to the right department, in ItalianRequest routingItalian36 of 36

How it is marked

The same instructions and material for every model, one answer a run, and a verdict that does not depend on how the answer sounds.

Pass or fail
An answer passes only if every fixed check passes. On the tasks that also carry a rubric, it has to meet most of its criteria too, and the criteria that matter most are mandatory. There is no partial credit.
Fixed checks
Structured answers are read field by field against the right values, with a tolerance on money and hours. Written answers are checked for what they must say, what they must not say and how long they may be. The same code marks every model.
The rubric grader
Rubric criteria are marked by gpt-5-6-sol at temperature zero, each with a reason. Every answer is marked twice, and a third mark settles any criterion the two disagree on. Its own verdict is ignored: code applies the pass rule to the criteria it marked. Before the run it was tested twice on answers whose verdict is known.
One protocol
Each run is one instruction message and one message with the material: no tools, no web search, no JSON mode, each model’s default sampling and up to 4,000 tokens of output, reasoning included. An answer cut off at that limit is marked as it stands. 1 run a model on each task.
Audited rules
Before an edition is published, every failure on a deciding rule is read. A rule that fails a correct answer for leaving out a detail the instructions never asked for is corrected, and those answers are marked again. A rule that catches a real mistake stays, however many models it fails.
Cost and time
A run costs what the provider billed for it, or, where the provider reports no cost, its tokens at the maker’s list price. Time runs from the request to the last word of the answer; calls that failed are left out of it.
Intervals
Every success rate carries a Wilson 95% interval. With a few dozen runs a model, a gap of several points between two models can be chance, and the bars say so.
The data
Tasks, checks and results are published under CC BY 4.0: cite AIGROW and link this page. Every task page carries the edition’s canary string, so the tasks can be kept out of training sets.

Put the models that pass to work

An AIGROW agent does this kind of work inside your own tools, against rules you approve in writing. To run the models yourself, API credit costs a quarter of list price.

October 2026 edition, pilot