
When should a clinical AI agent refuse to act?
A clinical agent can fail dangerously without saying anything false — by acting when the correct move was to gather a missing lab, abstain, or escalate to a clinician. I built a benchmark for that behaviour and a runtime guard that decides from policy and record data rather than from the model's proposal.
My work Case schema and four-way action space, fourteen synthetic cases, the data-driven policy and runtime guard, the evaluation harness, and a client-side port held to the original by a full-grid parity check.
Honest limit The cases and the policy share an author, so zero policy error shows self-consistency, not generalization. The agents are deterministic stand-ins, not language models.
Next Clinician-authored cases held out from the policy author, a second data seed, boundary traps at each threshold, and a real model-driven agent in place of the stand-ins.



