timniko.
Scene B · analytics · easy

What are customers searching for?

Turn on-site search logs into a purchasing decision.

opusmodelclaude-cli
100%verified3/3 checks
79%judgedby opus
20tool calls4 failed · 6 distinct
151swall clock20 turns
$0.84agent tokens · list pricecovered by the subscription

What the owner asked

Tell me what people are typing into my shop search.

Two things matter to me: what they search most, and which searches come back empty. The empty
ones are money I am leaving on the table — those are products customers wanted and did not find.

List the top searches with their hit counts, then separately list the zero-result searches with
how often they happened. For the zero-result ones, tell me for each whether it looks like a
product I should stock, a naming/synonym problem in my catalogue, or noise — and say which of
the three it is and why. End with a ranked shortlist of what to add, most valuable first.
A disciplined read-only run: 20 calls, four harmless failures (denied tables, one SQL quoting slip, one bad arg), no writes, and not a single unsupported number in the answer. The shop simply has no search data — one term, zero misses — and the model said so plainly, cross-checked blacklist/mapping/cache to rule out a configuration artefact, and quantified the underlying cause with real traffic figures. It loses points only where the scene's data makes the requested classification and product shortlist impossible; the substitute blocker ranking is sensible but is not what was asked.

Ground truth

Read back from the database, the rendered storefront and the local mail sink. No model involved.

ResultCheckWeight
PASS Shop state is unchanged (read-only scene)
no writes
3
PASS Called top_searches successfully
top_searches returned
3
PASS Answer separates zero-result searches from top searches
zero-result section present
2

Judgement

Scored by opus, which saw only the evidence bundle.

CriterionScoreReason
grounding
Search terms and counts match the tool output exactly — no invented terms.
10/10 Every figure traces to a call: term `klebrige`/1 search/2 hits/2026-07-30 19:06:56 from top_searches and tsuchanfrage (calls 1-3), the two matching articles 25/116 from the LIKE query (call 17), 103 visits/70 bots/33 human/11 IPs/1 conversion from visitor_statistics (call 14), and empty blacklist/mapping/erfolglos tables (calls 4, 19, 20) — nothing invented.
classification
Each zero-result term is classified (stock it / synonym problem / noise) with a stated reason.
7/10 No zero-result terms exist (tsuchanfrageerfolglos rowCount 0), and the answer states the stock/synonym/noise scheme cannot be applied and refuses to invent one — correct and honest, but it delivers no classification content, and it also verified the emptiness wasn't a blacklist/mapping artefact.
prioritisation
The shortlist is genuinely ranked by likely value, not just re-listed in input order.
6/10 There is no product shortlist to rank; instead it substitutes a sensibly ordered list of blockers (traffic first, then the lifetime-counter limitation, then 'several hundred searches' as a re-run threshold), which is useful but is not the ranked value list the owner asked for.
readability
Skimmable tables or lists; counts visible next to terms.
8/10 Numbered sections, a table with term/searches/hits/last-searched, bolded headline figures; loses a little for German-only prose and the near-empty table being unavoidably thin.
Highlight

It proved the zero result was real rather than a reporting artefact by separately checking tsuchanfrageblacklist, tsuchanfragemapping and tsuchanfragencache (calls 10, 19, 20) before declaring the gap list empty.

Weakness

Two of the four asked-for deliverables (classification, ranked product shortlist) are undeliverable from the data, so the owner receives an empty-data report — correct, but the substitute 'blocker' ranking only partly compensates.

How it worked

ToolCallsErrorsAvg
run_sql_read113297 ms
list_settings31299 ms
top_searches20313 ms
list_tables20341 ms
get_shop_info10276 ms
visitor_statistics10421 ms

The deliverable

# Shop-Suche: was Kunden eintippen

**Kurzfassung: Es gibt derzeit fast keine Suchdaten.** Der Shop hat in seiner gesamten
Historie genau **eine** protokollierte Suchanfrage und **null** ergebnislose Suchen.
Das ist kein Auswertungsproblem — die Zahlen sind so gespeichert.

## 1. Meistgesuchte Begriffe

| # | Suchbegriff | Suchen (gesamt) | Treffer | zuletzt gesucht |
|---|-------------|-----------------|---------|-----------------|
| 1 | `klebrige`  | 1               | 2       | 30.07.2026 19:06 |

Das war es. Weitere Begriffe sind nicht gespeichert.

Die beiden Treffer zu `klebrige` sind:
- **Klebrige Leinen-Bank** (Artikel 25)
- **Klebriger Leinen-Hut** (Artikel 116)

Die Suche selbst funktioniert also korrekt: Begriff eingegeben → passende Artikel geliefert →
Vorgang protokolliert. Die Protokollkette ist nachweislich intakt (Eintrag sowohl in der
Such-Statistik als auch im Such-Cache, gleicher Zeitstempel).

## 2. Suchen ohne Ergebnis

**Keine.** Die Tabelle der ergebnislosen Suchanfragen ist leer — 0 Einträge.

Damit gibt es aktuell **keinen einzigen belegten Nachfrage-Fehltreffer**. Zu jedem Begriff,
den je jemand gesucht hat, hat der Katalog geantwortet.

Zur Sicherheit geprüft, damit hier nichts künstlich leer aussieht:
- **Blacklist für Suchbegriffe: leer** — es wird nichts von der Auswertung ausgeschlossen.
- **Synonym-/Mapping-Tabelle: leer** — es werden keine Begriffe still umgeleitet, die sonst
  als "ohne Ergebnis" gezählt worden wären.

Beides bestätigt: Die Null ist echt, nicht wegkonfiguriert.

## 3. Einordnung der ergebnislosen Suchen

Sie wollten pro Fehlsuche eine Einstufung — Produkt fehlt / Benennungsproblem / Rauschen.
Diese Einstufung kann ich nicht liefern, weil es keine Fehlsuchen gibt. Eine Liste zu
erfinden wäre eine Liste erfundener Umsatzchancen.

Der einzige vorhandene Begriff `klebrige` ist im Übrigen **kein** Gap, sondern ein Treffer.

## 4. Ranked Shortlist: was aufnehmen?

Aus den Suchdaten lässt sich **keine** Produktempfehlung ableiten — eine einzige erfolgreiche
Suche trägt keine Rangliste. Statt einer Fantasie-Liste hier die Rangfolge dessen, was
tatsächlich zwischen Ihnen und dieser Auswertung steht:

1. **Traffic.** In den letzten 12 Monaten: **103 Besuche, davon 70 Bots** — also nur
   **33 menschliche Besuche** von **11 verschiedenen IP-Adressen**, 1 Bestellung.
   Bei dieser Größenordnung ist eine einzige Suchanfrage statistisch völlig plausibel.
   Ohne echte Besucher entsteht keine Suchstatistik, egal wie gut ausgewertet wird.
2. **Zeitfenster.** Die Such-Statistik zählt lebenslang, nicht pro Tag — es gibt keine
   Tageshistorie. Sobald Suchen anfallen, wachsen die Zähler, aber ein Trend
   ("diesen Monat vs. letzten") ist technisch nicht rekonstruierbar. Wenn Sie das brauchen,
   muss die Auswertung von Anfang an mitlaufen.
3. **Erneut auswerten.** Sinnvoll wird diese Analyse ab grob **einigen hundert Suchvorgängen**.
   Dann liefern die ergebnislosen Suchen genau das, was Sie suchen: bezahlte Nachfrage, auf
   die Ihr Katalog keine Antwort hat.

## Was ich verändert habe

Nichts. Dies war eine reine Auswertung — es wurden keine Einstellungen, Texte oder
Produktdaten angefasst.

What this costs to run

Setting it up
this benchmark run
$0.84

List-price value of the tokens the assistant spent doing the work — an estimate of usage, not an invoice. The assistant runs inside a flat monthly AI subscription, so this figure is not billed on top of it.

Running what it built
ongoing, per shop
not measured

Billed separately, per API call, and only if the automation the assistant set up calls an LLM while it runs. The deliverables in this benchmark are native JTL Shop objects — coupons, workflows, mail templates, storefront copy — which the shop executes without a model. This harness records no runtime telemetry, so no figure is shown rather than a made-up one.


Run 20260811-075829_b-demand-gaps_opus · shop reset to fixture before the run · restore with jtl restore 20260811-075829_b-demand-gaps_opus

© 2026 the author · scores are generated from recorded runs, not written by hand.