Tag: reliability
All the articles with the tag "reliability".
An Eval Suite That Never Blocks Is A Dashboard
The same 80% pass rate can mean a wording change or a working prompt injection. Severity tiers, baseline regressions, and an exit code CI respects.
Output Scoring Passes The Agent That Leaked Your Credentials
A correct answer reached by a forbidden path is a failure. Assert on the trajectory and the policy decisions, not the final message. Runnable suite.
Your Tool Definitions Are Prompts, And Most Are Bad Ones
An unbounded tool response costs 17x more tokens for the same answer. A runnable auditor for the four tool-design mistakes that make agents guess.
"I've Fixed The Failing Test" Is A Claim, Not A Completion Signal
Self-reported completion is not evidence. Independent deterministic verifiers, feedback re-injection, and why an LLM judging itself is theatre.