Skip to main navigation Skip to search Skip to main content

AI on Trial: LLM-as-a-Judge for Private and Reliable Clinical Decision-Making

Research output: Chapter in Book/Report/Conference proceedingConference contributionScientificpeer-review

Abstract

Artificial Intelligence (AI) is changing the way healthcare works by helping doctors make better decisions. Large Language Models (LLMs) play a key role in this shift, not only by answering medical questions, but also by actively collecting critical information through follow-up questions. However, effectively applying LLMs to complex Clinical Decision-Making (CDM) tasks remains a challenge. In this paper, we propose AI on Trial– an enhanced framework built on top of MDAgents. It assigns LLMs to work alone or in teams based on the complexity of medical queries. To improve the reliability of responses, we extend the original MDAgent system by incorporating LLM-as-a-Judge, a clinically informed evaluator that scores responses based on quality, completeness, safety and clarity, while also providing the reasoning behind its assessment. This judge also explains its reasoning using Chain-of-Thoughts (CoT’s). We also add a second judge that acts as a team of expert reviewers. It uses four LLM agents to carefully check the response and give a final score and confidence level. This two-step review process makes the system more accurate and trustworthy. Our results show that integrating the multi-agent evaluation with MDAgents improves performance, making it a strong tool for AI-based clinical support.
Original languageEnglish
Title of host publication2025 Annual Computer Security Applications Conference Workshops, ACSAC Workshops 2025
PublisherIEEE
Pages256-262
ISBN (Electronic)979-8-3315-4536-9
ISBN (Print)979-8-3315-4537-6
DOIs
Publication statusPublished - 2025
Publication typeA4 Article in conference proceedings
EventAnnual Computer Security Applications Conference Workshops - Honolulu, United States
Duration: 8 Dec 20259 Dec 2025

Conference

ConferenceAnnual Computer Security Applications Conference Workshops
Country/TerritoryUnited States
CityHonolulu
Period8/12/259/12/25

Publication forum classification

  • Publication forum level 1

Fingerprint

Dive into the research topics of 'AI on Trial: LLM-as-a-Judge for Private and Reliable Clinical Decision-Making'. Together they form a unique fingerprint.

Cite this