AI导师通过多智能体协作,实时指导放射科住院医师解读胸部X光片。
IMACT-CXR: An Interactive Multi-Agent Conversational Tutoring System for Chest X-Ray Interpretation
- 多智能体协同分析定位框、注视点与文字描述,实现交互式教学
- 相比基线模型,定位准确率和诊断推理能力显著提升
- 适合医学教育场景,支持真实DICOM数据与临床部署
IMACT-CXR 是一个基于 AutoGen 的多智能体对话式教学系统,用于辅助训练者解读胸部 X 光片。系统整合空间标注、注视分析、知识检索与图像引导推理,在单一工作流中实现多模态输入处理。该系统同时接收学习者的边界框、注视采样及自由文本观察,由专用智能体评估定位质量、生成苏格拉底式辅导、从 PubMed 检索证据、从 REFLACX 数据集推荐相似病例,并在掌握度不足或用户主动请求时调用 NV-Reason-CXR-3B 进行视觉-语言推理。贝叶斯知识追踪(BKT)动态维护各技能项的掌握度估计,驱动知识强化与病例相似性检索。基于 TensorFlow U-Net 的肺段分割模块提供解剖学感知的注视反馈,安全提示机制防止过早泄露真实标签。本文描述了系统架构、实现亮点及其与 REFLACX 数据集的集成,适用于真实 DICOM 病例。初步评估表明,IMACT-CXR 具有响应迅速、延迟可控、答案泄露可精准控制,并具备向实际住院医师培训系统扩展的潜力。相较于基线方法,其在定位精度与诊断推理能力上均有提升。
原文摘要 · Abstract (English)
IMACT-CXR is an interactive multi-agent conversational tutor that helps trainees interpret chest X-rays by unifying spatial annotation, gaze analysis, knowledge retrieval, and image-grounded reasoning in a single AutoGen-based workflow. The tutor simultaneously ingests learner bounding boxes, gaze samples, and free-text observations. Specialized agents evaluate localization quality, generate Socratic coaching, retrieve PubMed evidence, suggest similar cases from REFLACX, and trigger NV-Reason-CXR-3B for vision-language reasoning when mastery remains low or the learner explicitly asks. Bayesian Knowledge Tracing (BKT) maintains skill-specific mastery estimates that drive both knowledge reinforcement and case similarity retrieval. A lung-lobe segmentation module derived from a TensorFlow U-Net enables anatomically aware gaze feedback, and safety prompts prevent premature disclosure of ground-truth labels. We describe the system architecture, implementation highlights, and integration with the REFLACX dataset for real DICOM cases. IMACT-CXR demonstrates responsive tutoring flows with bounded latency, precise control over answer leakage, and extensibility toward live residency deployment. Preliminary evaluation shows improved localization and diagnostic reasoning compared to baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。