Policy-grounded
Every scenario connects to an explicit governing rule.
About
PiBench is an open research benchmark for evaluating whether explicit constraints remain effective while language-interactive agents communicate, use tools, and change state.
Purpose
An agent can understand a policy in text and still fail to enact it during a task. PiBench therefore treats the complete episode, rather than the final answer alone, as the unit of evaluation.
PiBench is a controlled evaluation instrument. It does not certify that a model, agent, or institution complies with law or organizational policy in deployment.
Recognition
PiBench received first place in the Agent Safety track of AgentX–AgentBeats Phase 1, organized through Berkeley RDI in conjunction with the Agentic AI MOOC.
Operating principles
Every scenario connects to an explicit governing rule.
Evaluation observes the surfaces on which the policy can be violated.
New policy packs use stable framework mechanisms rather than policy-specific core code.
Core team
AI engineer focused on building AI agents, evaluation harnesses, and policy-compliance benchmarks.
Machine Learning Engineer at Fiery, working on applied ML, large language models, and production AI systems.
AI and engineering leader with experience building ML infrastructure, voice AI, and agentic systems at Amazon, Samsung Research, and Vijil.
Open project
The codebase, current scenarios, evaluation reference, run reports, and literature survey are public. Research and dataset work remain active.