Best AI model for restaurants and bars
The tasks in the AIGROW small-business benchmark that a restaurant or bar would hand to AI, and how each of 36 models did on every one, marked pass or fail against written checks.
October 2026 edition, pilot, 1 run a model on each task.
- Qwen3.8 Flash
- 4 of 4
- Kimi K3
- 4 of 4
- GLM-5.3-Flash
- 4 of 4
- Claude Opus 5
- 4 of 4
- Claude Fable 5.1
- 4 of 4
- MiniMax M3
- 4 of 4
A tie goes to the better record across all 20 tasks.
The AIGROW small-business benchmark has 4 tasks a restaurant or bar would hand to AI: replying to negative reviews, triaging a business inbox, categorising bank transactions and categorising an Italian bank statement. In the October 2026 edition 6 models passed all 4; of those, GLM-5.3-Flash gave the cheapest passes across the benchmark, $0.831 per thousand.
Every model on these tasks
Most passes first; a tie goes to the better record across all 20 tasks. The cost of 1,000 passes is the whole benchmark’s, failed runs included.
| Model | Replying to negative reviews | Triaging a business inbox | Categorising bank transactions | Categorising an Italian bank statement | Passed | 1,000 passes |
|---|---|---|---|---|---|---|
| Qwen3.8 Flash | Pass | Pass | Pass | Pass | 4 of 4 | $1.09 |
| Kimi K3 | Pass | Pass | Pass | Pass | 4 of 4 | $18.44 |
| GLM-5.3-Flash | Pass | Pass | Pass | Pass | 4 of 4 | $0.831 |
| Claude Opus 5 | Pass | Pass | Pass | Pass | 4 of 4 | $16.89 |
| Claude Fable 5.1 | Pass | Pass | Pass | Pass | 4 of 4 | $33.11 |
| MiniMax M3 | Pass | Pass | Pass | Pass | 4 of 4 | $1.82 |
| Gemini 3.1 Pro (preview) | Fail | Pass | Pass | Pass | 3 of 4 | $29.67 |
| DeepSeek V4.1 Flash | Fail | Pass | Pass | Pass | 3 of 4 | $1.05 |
| DeepSeek V4 Pro | Pass | Pass | Fail | Pass | 3 of 4 | $5.14 |
| GLM-5.3 | Fail | Pass | Pass | Pass | 3 of 4 | $7.02 |
| Gemini 3.8 Flash | Fail | Pass | Pass | Pass | 3 of 4 | $7.17 |
| Grok 4.7 | Fail | Pass | Pass | Pass | 3 of 4 | $9.41 |
| Qwen3.7 Plus | Fail | Pass | Pass | Pass | 3 of 4 | $3.58 |
| Claude Sonnet 5 | Fail | Pass | Pass | Pass | 3 of 4 | $6.70 |
| GPT-5.6 Sol | Fail | Pass | Pass | Pass | 3 of 4 | $11.33 |
| Qwen3.8 Max | Fail | Pass | Pass | Pass | 3 of 4 | $16.32 |
| GPT-5.5 | Fail | Pass | Pass | Pass | 3 of 4 | $17.02 |
| GPT-6 Astra | Fail | Pass | Pass | Pass | 3 of 4 | $27.91 |
| GPT-5.6 Luna | Pass | Pass | Fail | Pass | 3 of 4 | $0.677 |
| Muse Glimmer 30B | Fail | Pass | Pass | Pass | 3 of 4 | $3.10 |
| GPT-5.6 Terra | Fail | Pass | Pass | Pass | 3 of 4 | $6.68 |
| Qwen3.8 27B | Fail | Pass | Pass | Pass | 3 of 4 | $7.30 |
| Claude Sonnet 4.6 | Fail | Pass | Pass | Pass | 3 of 4 | $11.24 |
| Kimi K2.6 | Fail | Pass | Pass | Pass | 3 of 4 | $11.68 |
| Grok 4.3 | Fail | Fail | Pass | Pass | 2 of 4 | $4.10 |
| gpt-oss-120b | Fail | Fail | Pass | Pass | 2 of 4 | $0.422 |
| Gemma 4 31B | Fail | Pass | Fail | Fail | 1 of 4 | $0.558 |
| Llama 4 Maverick | Fail | Pass | Fail | Fail | 1 of 4 | $0.665 |
| Gemini 3.1 Flash-Lite | Fail | Pass | Fail | Fail | 1 of 4 | $1.41 |
| GPT-5.4 mini | Fail | Pass | Fail | Fail | 1 of 4 | $3.35 |
| Mistral Medium 3.5 | Fail | Pass | Fail | Fail | 1 of 4 | $7.40 |
| Gemini 3.5 Flash-Lite | Fail | Fail | Fail | Pass | 1 of 4 | $2.47 |
| Claude Haiku 4.5 | Fail | Pass | Fail | Fail | 1 of 4 | $6.62 |
| GPT-5.4 nano | Fail | Fail | Fail | Fail | 0 of 4 | $1.51 |
| Ministral 3 14B | Fail | Fail | Fail | Fail | 0 of 4 | $1.07 |
| Mistral Small 4 | Fail | Fail | Fail | Fail | 0 of 4 | $1.33 |
The tasks
Each page shows the task as the models saw it, the traps in it, every check in words and what each model got wrong.
- Reply publicly to a negative restaurant review without revealing private details
- English. 8 of 36 runs passed.
- Triage a caterer's Monday inbox, including a phishing email and a bank-change request
- English. 30 of 36 runs passed.
- Categorise a café's monthly bank statement and total its operating costs
- English. 24 of 36 runs passed.
- Categorise an Italian bar's bank statement and total its movements
- Italian. 27 of 36 runs passed.
Hand this work to an agent
An AIGROW agent does work like this inside your own tools, against rules you approve in writing, and escalates what the rules do not cover instead of guessing.
October 2026 edition