A smart action is the Bench's recoverable gap turned into a move you can make today. It isn't a model guessing over raw rows — it's an agent that holds a question, a map of tools, and a reasoning loop, and fetches the answer from clean, pre-computed signals.
Like a sharp analyst dropped into an unfamiliar company with a really good filing system: they don't memorise the files — they know which drawer to open. In this snapshot it has already written 649 real smart actions across 17 stores.
Raw transactions never reach the agent. They're refined upstream into purpose-built gold marts — each one a pre-computed answer to a single question — then exposed as narrow, callable tools. Agents reason beautifully over clean signals and terribly over raw rows, so the Bench only ever hands them clean signals.
A real day at Independent · 04. The morning scan flags it's tracking soft, and the agent opens drawers in sequence until it can name the cause — then turns the recoverable gap into a play.
Independent · 04 currently has 0-0% RPMH and labor_cost entries in recent records (major blind spot).
Independent · 04 currently has 0-0% RPMH and labor_cost entries in recent records (major blind spot). Confidence high. Implement hourly reporting for total_hours and labor_cost_pct this_month so RPMH and cost controls can be calculated before next scheduling cycle. Expected impact: $100-$500 All recent entries show rpmh 0.0 and null labor_cost_pct; fixes enable meaningful scheduling.
These aren't illustrations. The system has generated 649 dated smart actions across 17 stores in this snapshot — each with a priority, a confidence, an expected dollar impact, and the full reasoning behind it.
Tuesdays at this location consistently run 30-36% below daily average revenue ($4394 vs $6550 overall) based on 13 observations with a tight range.
Tuesdays at this location consistently run 30-36% below daily average revenue ($4394 vs $6550 overall) based on 13 observations with a tight range. Confidence: high given repeated, recent samples. With Tue May 12 in next week's window, it's worth reviewing whether your current plans align with the expected lower volume that day. Expected impact: $1,700-$2,300 13 DOW observations; Tue avg revenue 4394 vs overall 6550.
Tuesdays at this location historically run about $2000-$2250 below the overall daily average (avg $4,281 vs.
Tuesdays at this location historically run about $2000-$2250 below the overall daily average (avg $4,281 vs. $6,409), based on 13 past Tuesdays with a tight range. As you finalize next week's plans, May 5 is worth a closer look given this pattern. It's worth reviewing whether your current plans align with the expected softer volume that day. Expected impact: $2,000-$2,250 Tuesday avg revenue $4,281.05 vs overall $6,409.22 (sample 13).
Tuesdays at this location historically run 30-37% below the daily average revenue ($4,066 vs $6,134) — based on 13 past Tuesdays with a tight range.
Tuesdays at this location historically run 30-37% below the daily average revenue ($4,066 vs $6,134) — based on 13 past Tuesdays with a tight range. High-confidence pattern. As you finalize next week's plan, Apr 28 is worth a closer look given this trend. It's worth reviewing whether your current plans align with the expected lower volume that day. Expected impact: $1,900-$2,200 Tuesday avg revenue $4,066 vs overall $6,134 (n=13).
Tuesdays at this location consistently run 20-28% below daily average revenue (roughly $5,145-$5,716 vs $7,145 overall).
Tuesdays at this location consistently run 20-28% below daily average revenue (roughly $5,145-$5,716 vs $7,145 overall). High-confidence pattern based on weekday history across the sample. With Tue May 26 in next week's window, it's worth reviewing whether your current plans align with the expected lower volume that day. Expected impact: $1,400-$2,000 Tuesday is the weakest weekday in the store's 88-day sample (average $5,190).
Operators don't open a BI tool — they ask. The agent answers conversationally on a phone, with the context baked in: every number carries its comparison, every trend its direction and size, and the date range it covers. Two ways it routes:
Scope is a property of the tools, not a rule we ask the model to follow. The agent reasons in display names, never raw IDs, and can only ever fetch what a user is allowed to see — which is exactly why this benchmark can be shown without exposing a single store's real name.
See the recoverable opportunity sized across the network, then dial in how often operators follow the plays.