When should a clinical AI agent refuse to act?
A clinical agent can fail dangerously without saying anything false — by acting when the right move was to gather a missing lab, abstain, or escalate. I built a benchmark and a runtime guard that sits outside the model, then ported the guard into a live sandbox.
Behind the build Four-way action space / data-driven policy / out-of-model enforcement / verified browser port
- 14
- synthetic clinical cases
- 0
- unsafe actions under enforcement
- 112
- episodes verified against the harness