Quick answers

AIGrow provides AI visibility monitoring, a business assistant, blog publishing and scoped automation services. Start with the free scan to inspect the output before paying.

Run the free scan

The scan is free and monitoring plans start at $29 a month. The operations audit is $290, the visibility audit is $490, and custom builds receive a written price before work starts.

See the plans

Run the free scan. It reads your site, asks several AI assistants what your customers ask, and points you at one thing. No signup required, and it will tell you if you do not need us.

Start now

Yes. Use the contact form for product, billing, support or project questions. Describe the goal and the system involved so the first reply can be specific.

Contact AIGrow
Benchmark · 4 tasks · October 2026 edition

Best AI model for shops

The tasks in the AIGROW small-business benchmark that a shop would hand to AI, and how each of 36 models did on every one, marked pass or fail against written checks.

October 2026 edition, pilot, 1 run a model on each task.

Most of the 4 tasks passed
4/412 models
Qwen3.8 Flash
4 of 4
Gemini 3.1 Pro (preview)
4 of 4
DeepSeek V4.1 Flash
4 of 4
DeepSeek V4 Pro
4 of 4
Gemini 3.8 Flash
4 of 4
Claude Sonnet 5
4 of 4

A tie goes to the better record across all 20 tasks.

The AIGROW small-business benchmark has 4 tasks a shop would hand to AI: answering refund requests, answering warranty claims in Italian, writing Italian product descriptions and writing meeting minutes. In the October 2026 edition 12 models passed all 4; of those, DeepSeek V4.1 Flash gave the cheapest passes across the benchmark, $1.05 per thousand.

Every model on these tasks

Most passes first; a tie goes to the better record across all 20 tasks. The cost of 1,000 passes is the whole benchmark’s, failed runs included.

Each model’s result on the tasks a shop hands out
ModelAnswering refund requestsAnswering warranty claims in ItalianWriting Italian product descriptionsWriting meeting minutesPassed1,000 passes
Qwen3.8 FlashPassPassPassPass4 of 4$1.09
Gemini 3.1 Pro (preview)PassPassPassPass4 of 4$29.67
DeepSeek V4.1 FlashPassPassPassPass4 of 4$1.05
DeepSeek V4 ProPassPassPassPass4 of 4$5.14
Gemini 3.8 FlashPassPassPassPass4 of 4$7.17
Claude Sonnet 5PassPassPassPass4 of 4$6.70
Qwen3.8 MaxPassPassPassPass4 of 4$16.32
Claude Opus 5PassPassPassPass4 of 4$16.89
GPT-5.5PassPassPassPass4 of 4$17.02
GPT-6 AstraPassPassPassPass4 of 4$27.91
GPT-5.6 TerraPassPassPassPass4 of 4$6.68
Kimi K2.6PassPassPassPass4 of 4$11.68
Kimi K3PassPassFailPass3 of 4$18.44
GLM-5.3-FlashPassPassFailPass3 of 4$0.831
GLM-5.3PassPassFailPass3 of 4$7.02
Grok 4.7PassPassFailPass3 of 4$9.41
Qwen3.7 PlusFailPassPassPass3 of 4$3.58
GPT-5.6 SolFailPassPassPass3 of 4$11.33
Claude Fable 5.1FailPassPassPass3 of 4$33.11
GPT-5.6 LunaFailPassPassPass3 of 4$0.677
MiniMax M3PassPassFailPass3 of 4$1.82
Claude Sonnet 4.6FailPassPassPass3 of 4$11.24
Gemma 4 31BFailPassPassPass3 of 4$0.558
Muse Glimmer 30BFailPassPassFail2 of 4$3.10
Grok 4.3FailPassFailPass2 of 4$4.10
Qwen3.8 27BFailPassFailPass2 of 4$7.30
gpt-oss-120bFailPassFailPass2 of 4$0.422
Llama 4 MaverickFailFailPassPass2 of 4$0.665
GPT-5.4 miniFailPassFailPass2 of 4$3.35
Mistral Medium 3.5FailPassFailPass2 of 4$7.40
Gemini 3.1 Flash-LiteFailFailPassFail1 of 4$1.41
Claude Haiku 4.5FailPassFailFail1 of 4$6.62
GPT-5.4 nanoFailFailFailPass1 of 4$1.51
Ministral 3 14BFailPassFailFail1 of 4$1.07
Gemini 3.5 Flash-LiteFailFailFailFail0 of 4$2.47
Mistral Small 4FailFailFailFail0 of 4$1.33

The tasks

Each page shows the task as the models saw it, the traps in it, every check in words and what each model got wrong.

Hand this work to an agent

An AIGROW agent does work like this inside your own tools, against rules you approve in writing, and escalates what the rules do not cover instead of guessing.

October 2026 edition