Study: The AI Judge You Pick Changes How Safe a Medical Chatbot Looks
Stress-testing Claude Opus 4.8, GPT-5.5, Grok 4.3 and Gemini 3.5 Flash on incomplete medical questions found LLM judges are more lenient than clinicians — and models score better when judged by their…