About

Policy compliance as an operational agent capability.

PiBench is an open research benchmark for evaluating whether explicit constraints remain effective while language-interactive agents communicate, use tools, and change state.

Purpose

Measure behavior, not policy recitation.

An agent can understand a policy in text and still fail to enact it during a task. PiBench therefore treats the complete episode, rather than the final answer alone, as the unit of evaluation.

PiBench is a controlled evaluation instrument. It does not certify that a model, agent, or institution complies with law or organizational policy in deployment.

Recognition

First place in the Agent Safety track.

PiBench received first place in the Agent Safety track of AgentX–AgentBeats Phase 1, organized through Berkeley RDI in conjunction with the Agentic AI MOOC.

Operating principles

Claims should remain auditable and versioned.

01

Policy-grounded

Every scenario connects to an explicit governing rule.

02

Episode-level

Evaluation observes the surfaces on which the policy can be violated.

03

Extensible by content

New policy packs use stable framework mechanisms rather than policy-specific core code.

Core team

PiBench research team.

Jyoti Ranjan Das, PiBench core team

Jyoti Ranjan Das

AI engineer focused on building AI agents, evaluation harnesses, and policy-compliance benchmarks.

Harshada Javeri, PiBench core team

Harshada Javeri

Machine Learning Engineer at Fiery, working on applied ML, large language models, and production AI systems.

Pradeep Das, PiBench core team

Pradeep Das

AI and engineering leader with experience building ML infrastructure, voice AI, and agentic systems at Amazon, Samsung Research, and Vijil.

Open project

Read the implementation and follow the work.

The codebase, current scenarios, evaluation reference, run reports, and literature survey are public. Research and dataset work remain active.