Model quality

How the agent is actually performing, measured rather than assumed.

Model quality in SigmaPointPi
/evals

The agent is tested continuously against suites covering the questions people actually ask: figures from the ledger, navigation, how-to answers and refusals.

A refusal suite matters as much as an accuracy suite. An agent that answers confidently when it should not know is worse than one that declines.

Where everything sits

Suites
Suites
Degraded
Degraded
Open regressions
Open regressions
Mean latest score
Mean latest score
Suite quality
Suite quality
Regression triage
Regression triage

How to work this page

Read accuracy by suite

Ledger questions, navigation, help desk and refusals score separately because they fail differently.

Watch the refusal rate

It should be non-zero. An agent that never declines is answering things it cannot know.

Check for regression

Scores are tracked across releases. A drop is investigated before it ships rather than after.

Read the failures

Individual failed cases are readable, which is how you judge whether a score means what it appears to.

On a phone

Model quality on iPhone 15 Pro Max

Every figure from the desktop appears here, stacked rather than reduced. Tables scroll inside themselves so the page never moves sideways, and figures keep their separators and their alignment at every width.

Questions people actually ask

Why show this at all?

Because trust is the thing people withhold, and a claim about accuracy without evidence is worth nothing.

What happens when a score drops?

It gates the release. The suites run in the build and a regression stops it.