Quick answers

AIGrow provides AI visibility monitoring, a business assistant, blog publishing and scoped automation services. Start with the free scan to inspect the output before paying.

Run the free scan

The scan is free and monitoring plans start at $29 a month. The operations audit is $290, the visibility audit is $490, and custom builds receive a written price before work starts.

See the plans

Run the free scan. It reads your site, asks several AI assistants what your customers ask, and points you at one thing. No signup required, and it will tell you if you do not need us.

Start now

Yes. Use the contact form for product, billing, support or project questions. Describe the goal and the system involved so the first reply can be specific.

Contact AIGrow
Benchmark · 6 tasks · October 2026 edition

Best AI model for accounting firms

The tasks in the AIGROW small-business benchmark that an accounting firm would hand to AI, and how each of 36 models did on every one, marked pass or fail against written checks.

October 2026 edition, pilot, 1 run a model on each task.

Most of the 6 tasks passed
6/615 models
Qwen3.8 Flash
6 of 6
Kimi K3
6 of 6
Gemini 3.1 Pro (preview)
6 of 6
GLM-5.3-Flash
6 of 6
DeepSeek V4.1 Flash
6 of 6
GLM-5.3
6 of 6

A tie goes to the better record across all 20 tasks.

The AIGROW small-business benchmark has 6 tasks an accounting firm would hand to AI: categorising an Italian bank statement, categorising bank transactions, extracting invoice data, extracting Italian supplier invoices, calculating Italian professional invoices and calculating overtime pay. In the October 2026 edition 15 models passed all 6; of those, gpt-oss-120b gave the cheapest passes across the benchmark, $0.422 per thousand.

Every model on these tasks

Most passes first; a tie goes to the better record across all 20 tasks. The cost of 1,000 passes is the whole benchmark’s, failed runs included.

Each model’s result on the tasks an accounting firm hands out
ModelCategorising an Italian bank statementCategorising bank transactionsExtracting invoice dataExtracting Italian supplier invoicesCalculating Italian professional invoicesCalculating overtime payPassed1,000 passes
Qwen3.8 FlashPassPassPassPassPassPass6 of 6$1.09
Kimi K3PassPassPassPassPassPass6 of 6$18.44
Gemini 3.1 Pro (preview)PassPassPassPassPassPass6 of 6$29.67
GLM-5.3-FlashPassPassPassPassPassPass6 of 6$0.831
DeepSeek V4.1 FlashPassPassPassPassPassPass6 of 6$1.05
GLM-5.3PassPassPassPassPassPass6 of 6$7.02
Grok 4.7PassPassPassPassPassPass6 of 6$9.41
Qwen3.7 PlusPassPassPassPassPassPass6 of 6$3.58
Claude Sonnet 5PassPassPassPassPassPass6 of 6$6.70
GPT-5.6 SolPassPassPassPassPassPass6 of 6$11.33
GPT-6 AstraPassPassPassPassPassPass6 of 6$27.91
Muse Glimmer 30BPassPassPassPassPassPass6 of 6$3.10
Grok 4.3PassPassPassPassPassPass6 of 6$4.10
Qwen3.8 27BPassPassPassPassPassPass6 of 6$7.30
gpt-oss-120bPassPassPassPassPassPass6 of 6$0.422
DeepSeek V4 ProPassFailPassPassPassPass5 of 6$5.14
Gemini 3.8 FlashPassPassFailPassPassPass5 of 6$7.17
Qwen3.8 MaxPassPassFailPassPassPass5 of 6$16.32
Claude Opus 5PassPassPassPassPassFail5 of 6$16.89
GPT-5.5PassPassPassPassPassFail5 of 6$17.02
Claude Fable 5.1PassPassPassPassPassFail5 of 6$33.11
GPT-5.6 TerraPassPassPassPassPassFail5 of 6$6.68
MiniMax M3PassPassPassFailPassPass5 of 6$1.82
Claude Sonnet 4.6PassPassPassPassPassFail5 of 6$11.24
Kimi K2.6PassPassPassPassPassFail5 of 6$11.68
GPT-5.6 LunaPassFailPassPassPassFail4 of 6$0.677
Gemini 3.5 Flash-LitePassFailPassPassPassFail4 of 6$2.47
Gemma 4 31BFailFailPassPassPassFail3 of 6$0.558
Gemini 3.1 Flash-LiteFailFailPassPassPassFail3 of 6$1.41
GPT-5.4 miniFailFailPassPassPassFail3 of 6$3.35
Claude Haiku 4.5FailFailPassPassPassFail3 of 6$6.62
GPT-5.4 nanoFailFailPassPassPassFail3 of 6$1.51
Llama 4 MaverickFailFailPassFailPassFail2 of 6$0.665
Mistral Medium 3.5FailFailPassFailPassFail2 of 6$7.40
Ministral 3 14BFailFailPassFailPassFail2 of 6$1.07
Mistral Small 4FailFailPassFailPassFail2 of 6$1.33

The tasks

Each page shows the task as the models saw it, the traps in it, every check in words and what each model got wrong.

Hand this work to an agent

An AIGROW agent does work like this inside your own tools, against rules you approve in writing, and escalates what the rules do not cover instead of guessing.

October 2026 edition