Model quality
How the agent is actually performing, measured rather than assumed.

The agent is tested continuously against suites covering the questions people actually ask: figures from the ledger, navigation, how-to answers and refusals.
A refusal suite matters as much as an accuracy suite. An agent that answers confidently when it should not know is worse than one that declines.
Where everything sits






How to work this page
Ledger questions, navigation, help desk and refusals score separately because they fail differently.
It should be non-zero. An agent that never declines is answering things it cannot know.
Scores are tracked across releases. A drop is investigated before it ships rather than after.
Individual failed cases are readable, which is how you judge whether a score means what it appears to.
On a phone

Every figure from the desktop appears here, stacked rather than reduced. Tables scroll inside themselves so the page never moves sideways, and figures keep their separators and their alignment at every width.
Questions people actually ask
Because trust is the thing people withhold, and a claim about accuracy without evidence is worth nothing.
It gates the release. The suites run in the build and a regression stops it.