Best AI model for shops
The tasks in the AIGROW small-business benchmark that a shop would hand to AI, and how each of 36 models did on every one, marked pass or fail against written checks.
October 2026 edition, pilot, 1 run a model on each task.
- Qwen3.8 Flash
- 4 of 4
- Gemini 3.1 Pro (preview)
- 4 of 4
- DeepSeek V4.1 Flash
- 4 of 4
- DeepSeek V4 Pro
- 4 of 4
- Gemini 3.8 Flash
- 4 of 4
- Claude Sonnet 5
- 4 of 4
A tie goes to the better record across all 20 tasks.
The AIGROW small-business benchmark has 4 tasks a shop would hand to AI: answering refund requests, answering warranty claims in Italian, writing Italian product descriptions and writing meeting minutes. In the October 2026 edition 12 models passed all 4; of those, DeepSeek V4.1 Flash gave the cheapest passes across the benchmark, $1.05 per thousand.
Every model on these tasks
Most passes first; a tie goes to the better record across all 20 tasks. The cost of 1,000 passes is the whole benchmark’s, failed runs included.
| Model | Answering refund requests | Answering warranty claims in Italian | Writing Italian product descriptions | Writing meeting minutes | Passed | 1,000 passes |
|---|---|---|---|---|---|---|
| Qwen3.8 Flash | Pass | Pass | Pass | Pass | 4 of 4 | $1.09 |
| Gemini 3.1 Pro (preview) | Pass | Pass | Pass | Pass | 4 of 4 | $29.67 |
| DeepSeek V4.1 Flash | Pass | Pass | Pass | Pass | 4 of 4 | $1.05 |
| DeepSeek V4 Pro | Pass | Pass | Pass | Pass | 4 of 4 | $5.14 |
| Gemini 3.8 Flash | Pass | Pass | Pass | Pass | 4 of 4 | $7.17 |
| Claude Sonnet 5 | Pass | Pass | Pass | Pass | 4 of 4 | $6.70 |
| Qwen3.8 Max | Pass | Pass | Pass | Pass | 4 of 4 | $16.32 |
| Claude Opus 5 | Pass | Pass | Pass | Pass | 4 of 4 | $16.89 |
| GPT-5.5 | Pass | Pass | Pass | Pass | 4 of 4 | $17.02 |
| GPT-6 Astra | Pass | Pass | Pass | Pass | 4 of 4 | $27.91 |
| GPT-5.6 Terra | Pass | Pass | Pass | Pass | 4 of 4 | $6.68 |
| Kimi K2.6 | Pass | Pass | Pass | Pass | 4 of 4 | $11.68 |
| Kimi K3 | Pass | Pass | Fail | Pass | 3 of 4 | $18.44 |
| GLM-5.3-Flash | Pass | Pass | Fail | Pass | 3 of 4 | $0.831 |
| GLM-5.3 | Pass | Pass | Fail | Pass | 3 of 4 | $7.02 |
| Grok 4.7 | Pass | Pass | Fail | Pass | 3 of 4 | $9.41 |
| Qwen3.7 Plus | Fail | Pass | Pass | Pass | 3 of 4 | $3.58 |
| GPT-5.6 Sol | Fail | Pass | Pass | Pass | 3 of 4 | $11.33 |
| Claude Fable 5.1 | Fail | Pass | Pass | Pass | 3 of 4 | $33.11 |
| GPT-5.6 Luna | Fail | Pass | Pass | Pass | 3 of 4 | $0.677 |
| MiniMax M3 | Pass | Pass | Fail | Pass | 3 of 4 | $1.82 |
| Claude Sonnet 4.6 | Fail | Pass | Pass | Pass | 3 of 4 | $11.24 |
| Gemma 4 31B | Fail | Pass | Pass | Pass | 3 of 4 | $0.558 |
| Muse Glimmer 30B | Fail | Pass | Pass | Fail | 2 of 4 | $3.10 |
| Grok 4.3 | Fail | Pass | Fail | Pass | 2 of 4 | $4.10 |
| Qwen3.8 27B | Fail | Pass | Fail | Pass | 2 of 4 | $7.30 |
| gpt-oss-120b | Fail | Pass | Fail | Pass | 2 of 4 | $0.422 |
| Llama 4 Maverick | Fail | Fail | Pass | Pass | 2 of 4 | $0.665 |
| GPT-5.4 mini | Fail | Pass | Fail | Pass | 2 of 4 | $3.35 |
| Mistral Medium 3.5 | Fail | Pass | Fail | Pass | 2 of 4 | $7.40 |
| Gemini 3.1 Flash-Lite | Fail | Fail | Pass | Fail | 1 of 4 | $1.41 |
| Claude Haiku 4.5 | Fail | Pass | Fail | Fail | 1 of 4 | $6.62 |
| GPT-5.4 nano | Fail | Fail | Fail | Pass | 1 of 4 | $1.51 |
| Ministral 3 14B | Fail | Pass | Fail | Fail | 1 of 4 | $1.07 |
| Gemini 3.5 Flash-Lite | Fail | Fail | Fail | Fail | 0 of 4 | $2.47 |
| Mistral Small 4 | Fail | Fail | Fail | Fail | 0 of 4 | $1.33 |
The tasks
Each page shows the task as the models saw it, the traps in it, every check in words and what each model got wrong.
- Reply to a customer asking for a refund outside the refund window
- English. 17 of 36 runs passed.
- Reply in Italian to a customer who wrongly thinks her warranty has expired
- Italian. 31 of 36 runs passed.
- Write an Italian shop description for an olive oil that respects labelling rules
- Italian. 21 of 36 runs passed.
- Turn a bakery team meeting into decisions, owners and due dates
- English. 30 of 36 runs passed.
Hand this work to an agent
An AIGROW agent does work like this inside your own tools, against rules you approve in writing, and escalates what the rules do not cover instead of guessing.
October 2026 edition