Your Evaluation Sandbox Holds Production Authority
An evaluation environment can be disposable while its credentials are not. A practical design for separating test workloads, provider keys, and production authority.
notebook / tag
9 entries with this tag.
An evaluation environment can be disposable while its credentials are not. A practical design for separating test workloads, provider keys, and production authority.
Aur0ra operators reportedly persuaded a coding agent that real intrusions were authorized tests. The lesson is not simply that models can be fooled. Authorization has to exist outside the conversation.
ToolHazard turns indirect prompt-injection testing into executable, stateful evaluation. The useful lesson is not its leaderboard. It is how to make agent security a repeatable release gate.
Anthropic found three real intrusions inside cyber evaluations whose prompts claimed the internet was unavailable. A safe range needs machine-enforced scope, verified egress, and live boundary detection.
An audit of Aria's autonomous loop found hundreds of goals, almost no progress, and a completion system that rewarded plausible output.
A self-observation feature for Aria showed why metrics should remain available without becoming permanent instructions.
What two rounds of testing Aria's language and image models taught me about speed, benchmarks, and knowing when a score is wrong.
A quick way to tell whether your agent setup is ready to grow or still held together by one-off fixes.
Four simple work patterns for agents that use tools, make changes, test the result, and recover from interruptions.