Judge-Dependent Safety Gains and Model-Specific Helpfulness Costs of Evidence-Sufficiency Prompting in Clinical LLMs
A prompting trick meant to stop clinical AI from overconfident answers looks safer or useless depending entirely on which AI judges it.