Abstract
Artificial Intelligence (AI) is changing the way healthcare works by helping doctors make better decisions. Large Language Models (LLMs) play a key role in this shift, not only by answering medical questions, but also by actively collecting critical information through follow-up questions. However, effectively applying LLMs to complex Clinical Decision-Making (CDM) tasks remains a challenge. In this paper, we propose AI on Trial– an enhanced framework built on top of MDAgents. It assigns LLMs to work alone or in teams based on the complexity of medical queries. To improve the reliability of responses, we extend the original MDAgent system by incorporating LLM-as-a-Judge, a clinically informed evaluator that scores responses based on quality, completeness, safety and clarity, while also providing the reasoning behind its assessment. This judge also explains its reasoning using Chain-of-Thoughts (CoT’s). We also add a second judge that acts as a team of expert reviewers. It uses four LLM agents to carefully check the response and give a final score and confidence level. This two-step review process makes the system more accurate and trustworthy. Our results show that integrating the multi-agent evaluation with MDAgents improves performance, making it a strong tool for AI-based clinical support.
| Original language | English |
|---|---|
| Title of host publication | 2025 Annual Computer Security Applications Conference Workshops, ACSAC Workshops 2025 |
| Publisher | IEEE |
| Pages | 256-262 |
| ISBN (Electronic) | 979-8-3315-4536-9 |
| ISBN (Print) | 979-8-3315-4537-6 |
| DOIs | |
| Publication status | Published - 2025 |
| Publication type | A4 Article in conference proceedings |
| Event | Annual Computer Security Applications Conference Workshops - Honolulu, United States Duration: 8 Dec 2025 → 9 Dec 2025 |
Conference
| Conference | Annual Computer Security Applications Conference Workshops |
|---|---|
| Country/Territory | United States |
| City | Honolulu |
| Period | 8/12/25 → 9/12/25 |
Publication forum classification
- Publication forum level 1
Fingerprint
Dive into the research topics of 'AI on Trial: LLM-as-a-Judge for Private and Reliable Clinical Decision-Making'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver