
An Eval Suite That Never Blocks Is A Dashboard
The same 80% pass rate can mean a wording change or a working prompt injection. Severity tiers, baseline regressions, and an exit code CI respects.
All the articles with the tag "Reliability".

The same 80% pass rate can mean a wording change or a working prompt injection. Severity tiers, baseline regressions, and an exit code CI respects.

A correct answer reached by a forbidden path is a failure. Assert on the trajectory and the policy decisions, not the final message. Runnable suite.

An unbounded tool response costs 17x more tokens for the same answer. A runnable auditor for the four tool-design mistakes that make agents guess.

Self-reported completion is not evidence. Independent deterministic verifiers, feedback re-injection, and why an LLM judging itself is theatre.

At-least-once execution is the default everywhere. Checkpoints make resumption cheap; idempotency keys make it safe. Runnable code and the bug I hit.

Dropping the oldest messages is amnesia with a token budget. Compact by turn group, pin the goal and the denials, and never orphan a tool call.