针对眼科大模型幻觉问题,构建可追溯的推理框架提升诊断准确性。
EH-Benchmark Ophthalmic Hallucination Benchmark and Agent-Driven Top-Down Traceable Reasoning Workflow
- 分视觉理解与逻辑组合两类设计幻觉评估体系
- 多智能体三阶段流程降低幻觉率,提升诊断可靠性
- 适合医疗AI研发者和临床辅助系统开发者参考
医学大语言模型(MLLMs)在眼科诊断中具有重要价值,但其准确率受限于眼科知识不足、视觉定位与推理能力弱以及多模态眼科数据稀缺,导致病变检测与疾病诊断精度下降。现有医疗基准无法有效评估各类幻觉或提供缓解方案。为此,我们提出EH-Benchmark,一个专为眼科幻觉评估设计的新基准,将幻觉按任务与错误类型分为视觉理解与逻辑组合两大类,每类包含多个子类。鉴于MLLM主要依赖语言推理而非视觉处理,我们提出以智能体为中心的三阶段框架:知识层检索、任务级案例研究与结果层验证。实验表明,该多智能体框架显著降低两类幻觉,提升准确率、可解释性与可靠性。项目开源地址:https://github.com/ppxy1/EH-Benchmark。
原文摘要 · Abstract (English)
Medical Large Language Models (MLLMs) play a crucial role in ophthalmic diagnosis, holding significant potential to address vision-threatening diseases. However, their accuracy is constrained by hallucinations stemming from limited ophthalmic knowledge, insufficient visual localization and reasoning capabilities, and a scarcity of multimodal ophthalmic data, which collectively impede precise lesion detection and disease diagnosis. Furthermore, existing medical benchmarks fail to effectively evaluate various types of hallucinations or provide actionable solutions to mitigate them. To address the above challenges, we introduce EH-Benchmark, a novel ophthalmology benchmark designed to evaluate hallucinations in MLLMs. We categorize MLLMs' hallucinations based on specific tasks and error types into two primary classes: Visual Understanding and Logical Composition, each comprising multiple subclasses. Given that MLLMs predominantly rely on language-based reasoning rather than visual processing, we propose an agent-centric, three-phase framework, including the Knowledge-Level Retrieval stage, the Task-Level Case Studies stage, and the Result-Level Validation stage. Experimental results show that our multi-agent framework significantly mitigates both types of hallucinations, enhancing accuracy, interpretability, and reliability. Our project is available at https://github.com/ppxy1/EH-Benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。