Pick a scenario and use a real AI app. Beside it, the bench lights up every step it takes, the exact prompt the model was given, what came back, and what it cost.
Each one is a real AI app, running from recordings of real model calls. Start one, use the app, and watch the bench beside it. Open the lab →
Readable at a glance. Every number underneath is real.
Every step the app could take. The path it actually took lights up; the branches it skipped stay dashed.
What the model was told, word for word, and exactly what it said back.
How long each step took, side by side, including time spent waiting for a person.
Tokens in and out, and what every model call cost.
An app reports what it does as it runs, and registers a map of every step it could take. The bench draws the map, then fills it in live as events arrive, or replays a recording with its original timing.
It sits on top of the tracing an app already has, OpenTelemetry or its own events, so it works with whatever you use to observe AI in production. It never changes the app; approvals and every other decision stay where they belong, in the app.
Apps can also bring their own story: custom panels that explain their runs, the way comments explain code.