Every score below comes from a run against the same reset shop: half of it read back from the database and the rendered storefront, half of it judged from the evidence by a separate model.
Each card opens the best-scoring run for that scene — the full evidence bundle: what the shop looked like before, every tool call, what was read back out of the database afterwards.
Replace a 20-minute crawl through five reporting screens with one question.
Turn on-site search logs into a purchasing decision.
Turn raw log noise into a triaged, plain-language incident list.
Event-driven automation without leaving the shop or writing code.
Two systems — coupons and workflows — chained into one campaign.
Backend action plus visible storefront change — the full campaign loop.
Retention automation on a trigger most owners never find in the UI.
Creative seasonal work with the same hard requirements as Black Friday.
| Model | Verified | Judged | Calls | Errors | Avg time | ||
|---|---|---|---|---|---|---|---|
| opus 10 scenes |
100% | 85% | 250 | 15 | 218s | ||
| gemini-flash 10 scenes |
100% | 78% | 247 | 10 | 87s | ||
| gemini-pro-3.1 10 scenes |
100% | 71% | 104 | 9 | 983s | ||
| sonnet 10 scenes |
95% | 76% | 152 | 13 | 140s | ||
| gemini-flash-3.5 10 scenes |
100% | 71% | 237 | 7 | 19298s | ||
| haiku 10 scenes |
87% | 59% | 169 | 19 | 159s | ||
Verified percentage on top, the judge's score underneath, and a link to the run itself in the row below it.
| Scene | sonnet | opus | haiku | gemini-pro-3.1 | gemini-flash | gemini-flash-3.5 | |
|---|---|---|---|---|---|---|---|
| How is my shop doing? analytics · easy | 100% judge 85% | 100% judge 96% | 100% judge 77% | 100% judge 81% | 100% judge 83% | 100% judge 75% | |
| open the full run | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | |
| What are customers searching for? analytics · easy | 100% judge 79% | 100% judge 79% | 100% judge 63% | 100% judge 62% | 100% judge 70% | 100% judge 57% | |
| open the full run | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | |
| Check my logs for problems operations · easy | 100% judge 85% | 100% judge 88% | 100% judge 58% | 100% judge 74% | 100% judge 77% | 100% judge 69% | |
| open the full run | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | |
| Create a summer discount admin · medium | 100% judge 92% | 100% judge 94% | 100% judge 80% | 100% judge 82% | 100% judge 92% | 100% judge 90% | |
| open the full run | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | |
| Send a thank-you email on every order workflow · medium | 100% judge 79% | 100% judge 88% | 100% judge 67% | 100% judge 81% | 100% judge 83% | 100% judge 79% | |
| open the full run | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | |
| Welcome new customers with a discount workflow + admin · medium | 100% judge 84% | 100% judge 83% | 100% judge 75% | 100% judge 68% | 100% judge 81% | 100% judge 77% | |
| open the full run | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | |
| Launch a Black Friday campaign campaign · hard | 71% judge 31% | 100% judge 73% | 71% judge 34% | 100% judge 74% | 100% judge 78% | 100% judge 70% | |
| open the full run | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | |
| Convert guest checkouts workflow · medium | 100% judge 82% | 100% judge 83% | 0% judge 39% | 100% judge 80% | 100% judge 82% | 100% judge 71% | |
| open the full run | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | |
| Prepare for Christmas campaign · hard | 100% judge 68% | 100% judge 75% | 100% judge 43% | 100% judge 54% | 100% judge 75% | 100% judge 64% | |
| open the full run | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | |
| Give me strategic advice consulting · medium | 75% judge 77% | 100% judge 86% | 100% judge 54% | 100% judge 55% | 100% judge 57% | 100% judge 55% | |
| open the full run | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ | Report ↗ |
The assistant doing the work runs inside a flat monthly AI subscription — its own token usage is not invoiced per run, so no per-run price is quoted here. What a shop would pay separately is the runtime LLM cost of the automation the assistant sets up, and only if that automation calls a model while it runs; the deliverables in this benchmark are native JTL Shop objects that run without one. Each run page shows both halves, and says plainly that the second is not measured here.
60 runs · generated by jtl report
© 2026 the author · scores are generated from recorded runs, not written by hand.