When should a clinical AI agent refuse to act?
A clinical agent can fail dangerously without saying anything false — by acting when the right move was to gather a missing lab, abstain, or escalate. I built a benchmark and a runtime guard that sits outside the model, with a 20-case model evaluation and an earlier interactive browser demo.
Latest evaluation Guard off → on, on synthetic FHIR cases. Small-benchmark results, not clinical validation. The illustration and sandbox show the earlier 14-case deterministic demo.
- 20
- synthetic evaluation cases
- 0.35 → 0
- 7B model unsafe-action rate
- 0.375 → 0
- frontier model over-refusal rate