Best AI model for trades and services
The tasks in the AIGROW small-business benchmark that a trade or service business would hand to AI, and how each of 36 models did on every one, marked pass or fail against written checks.
October 2026 edition, pilot, 1 run a model on each task.
- Kimi K3
- 6 of 6
- Gemini 3.1 Pro (preview)
- 6 of 6
- GLM-5.3
- 6 of 6
- Gemini 3.8 Flash
- 6 of 6
- Grok 4.7
- 6 of 6
- Grok 4.3
- 6 of 6
A tie goes to the better record across all 20 tasks.
The AIGROW small-business benchmark has 6 tasks a trade or service business would hand to AI. In the October 2026 edition 8 models passed all 6; of those, gpt-oss-120b gave the cheapest passes across the benchmark, $0.422 per thousand. Each task page shows the traps and every model’s mistakes.
Every model on these tasks
Most passes first; a tie goes to the better record across all 20 tasks. The cost of 1,000 passes is the whole benchmark’s, failed runs included.
| Model | Pricing a job from a price list | Writing quote emails | Scoring sales leads | Routing customer messages | Calculating overtime pay | Answering staff handbook questions | Passed | 1,000 passes |
|---|---|---|---|---|---|---|---|---|
| Kimi K3 | Pass | Pass | Pass | Pass | Pass | Pass | 6 of 6 | $18.44 |
| Gemini 3.1 Pro (preview) | Pass | Pass | Pass | Pass | Pass | Pass | 6 of 6 | $29.67 |
| GLM-5.3 | Pass | Pass | Pass | Pass | Pass | Pass | 6 of 6 | $7.02 |
| Gemini 3.8 Flash | Pass | Pass | Pass | Pass | Pass | Pass | 6 of 6 | $7.17 |
| Grok 4.7 | Pass | Pass | Pass | Pass | Pass | Pass | 6 of 6 | $9.41 |
| Grok 4.3 | Pass | Pass | Pass | Pass | Pass | Pass | 6 of 6 | $4.10 |
| Qwen3.8 27B | Pass | Pass | Pass | Pass | Pass | Pass | 6 of 6 | $7.30 |
| gpt-oss-120b | Pass | Pass | Pass | Pass | Pass | Pass | 6 of 6 | $0.422 |
| Qwen3.8 Flash | Fail | Pass | Pass | Pass | Pass | Pass | 5 of 6 | $1.09 |
| GLM-5.3-Flash | Pass | Fail | Pass | Pass | Pass | Pass | 5 of 6 | $0.831 |
| DeepSeek V4.1 Flash | Pass | Fail | Pass | Pass | Pass | Pass | 5 of 6 | $1.05 |
| DeepSeek V4 Pro | Pass | Fail | Pass | Pass | Pass | Pass | 5 of 6 | $5.14 |
| Qwen3.7 Plus | Pass | Fail | Pass | Pass | Pass | Pass | 5 of 6 | $3.58 |
| GPT-5.6 Sol | Pass | Fail | Pass | Pass | Pass | Pass | 5 of 6 | $11.33 |
| Qwen3.8 Max | Pass | Fail | Pass | Pass | Pass | Pass | 5 of 6 | $16.32 |
| GPT-6 Astra | Pass | Fail | Pass | Pass | Pass | Pass | 5 of 6 | $27.91 |
| Muse Glimmer 30B | Pass | Fail | Pass | Pass | Pass | Pass | 5 of 6 | $3.10 |
| Claude Sonnet 5 | Fail | Fail | Pass | Pass | Pass | Pass | 4 of 6 | $6.70 |
| GPT-5.5 | Pass | Fail | Pass | Pass | Fail | Pass | 4 of 6 | $17.02 |
| Claude Fable 5.1 | Pass | Fail | Pass | Pass | Fail | Pass | 4 of 6 | $33.11 |
| GPT-5.6 Luna | Pass | Fail | Pass | Pass | Fail | Pass | 4 of 6 | $0.677 |
| MiniMax M3 | Fail | Pass | Pass | Pass | Pass | Fail | 4 of 6 | $1.82 |
| Claude Opus 5 | Fail | Fail | Pass | Pass | Fail | Pass | 3 of 6 | $16.89 |
| GPT-5.6 Terra | Fail | Fail | Pass | Pass | Fail | Pass | 3 of 6 | $6.68 |
| Claude Sonnet 4.6 | Fail | Fail | Pass | Pass | Fail | Pass | 3 of 6 | $11.24 |
| Kimi K2.6 | Fail | Pass | Fail | Pass | Fail | Pass | 3 of 6 | $11.68 |
| Gemma 4 31B | Pass | Fail | Pass | Pass | Fail | Fail | 3 of 6 | $0.558 |
| Gemini 3.1 Flash-Lite | Pass | Pass | Fail | Pass | Fail | Fail | 3 of 6 | $1.41 |
| Llama 4 Maverick | Fail | Pass | Fail | Pass | Fail | Fail | 2 of 6 | $0.665 |
| Mistral Medium 3.5 | Fail | Pass | Fail | Pass | Fail | Fail | 2 of 6 | $7.40 |
| Gemini 3.5 Flash-Lite | Fail | Pass | Fail | Pass | Fail | Fail | 2 of 6 | $2.47 |
| Claude Haiku 4.5 | Fail | Fail | Fail | Pass | Fail | Fail | 1 of 6 | $6.62 |
| GPT-5.4 nano | Fail | Pass | Fail | Fail | Fail | Fail | 1 of 6 | $1.51 |
| Mistral Small 4 | Fail | Fail | Fail | Pass | Fail | Fail | 1 of 6 | $1.33 |
| GPT-5.4 mini | Fail | Fail | Fail | Fail | Fail | Fail | 0 of 6 | $3.35 |
| Ministral 3 14B | Fail | Fail | Fail | Fail | Fail | Fail | 0 of 6 | $1.07 |
The tasks
Each page shows the task as the models saw it, the traps in it, every check in words and what each model got wrong.
- Price a decorating job from survey measurements and the price list
- English. 21 of 36 runs passed.
- Email a garden landscaping quote that handles a VAT request and a protected tree
- English. 16 of 36 runs passed.
- Score five inbound cleaning leads against the sales team's rules
- English. 26 of 36 runs passed.
- Route ten tenant messages to the right team with the right priority
- English. 33 of 36 runs passed.
- Work out a cleaner's weekly gross pay with breaks, overtime and Sunday rates
- English. 19 of 36 runs passed.
- Answer a part-time employee's leave questions from the staff handbook
- English. 25 of 36 runs passed.
Hand this work to an agent
An AIGROW agent does work like this inside your own tools, against rules you approve in writing, and escalates what the rules do not cover instead of guessing.
October 2026 edition