arXiv:2608.30352cs.AIcs.HC2026-08

用专家注视和口述数据训练视觉与文档双引导系统,提升眼底OCT诊断效率。

Co-Annotator: Expert-Distilled ViT and VLM for Visual and Documentation Guidance in Age-Related Macular Degeneration

论文配图:Co-Annotator: Expert-Distilled ViT and VLM for Visual and Documentation Guidance in Age-Related Macular Degeneration
图 1 · 摘自论文原文
  • 通过专家注视数据训练视觉变压器,生成关注区域;结合语义约束的视觉语言模型预填病灶摘要。
  • 联合使用时诊断速度提升40%,报告编辑时间减少67%,且不降低诊断准确率。
  • 适合眼科医生、临床研究者及医学AI开发者,特别适用于影像辅助诊断场景。

临床AI通常只优化预测性能,而不考虑医生如何决定观察位置和书写内容。本文提出Co-Annotator,将专家注视和口述信息提炼为两个指导模块:一个对齐注视点的视觉变换器(ViT),用于生成注视对齐的关注区域(AOIs);一个基于本体的视觉语言模型(VLM),可为视网膜光学相干断层扫描(OCT)图像预填可编辑的生物标志物摘要。首先,我们采集了专家注视和口述数据(US1)以训练模型,显著提升了诊断准确率和病灶生成效果。随后,在眼科住院医师中进行对照研究(US2),证实两种模态均安全且独立有益:AOI引导带来持久的感知效率提升,而VLM引导使病灶文档覆盖面翻倍以上。在两家学术机构的联合部署中(US3),同时提供两种引导时,效率提升远超单一模态:每分钟正确诊断数增加40%,报告编辑时间减少67%,且未影响诊断准确性。值得注意的是,在US2中单个模态并未提升引导期间的效率,因此在US3中联合引导带来的效率增益尤为突出。专家提炼的多模态引导能同时缓解视觉搜索负担和文档书写压力,且不牺牲现有诊断精度。

原文摘要 · Abstract (English)

Clinical AI often optimizes predictive performance without engaging how clinicians decide where to look and what to write. We present Co-Annotator, which distills expert gaze and dictation into two guidance components: a gaze-aligned Vision Transformer producing fixation-aligned areas of interest (AOIs), and an ontology-bounded vision-language model (VLM) that pre-fills editable biomarker summaries for retinal optical coherence tomography (OCT). We first collect expert gaze and dictations (US1) to train the models, significantly improving diagnostic accuracy and biomarker generation. We then deploy the system with ophthalmology residents: a controlled resident study (US2) confirmed each modality is safe and independently beneficial, with AOI guidance producing lasting perceptual efficiency gains through post-guidance carryover and VLM guidance more than doubling biomarker documentation breadth. In a combined deployment across two academic institutions (US3), providing both modalities simultaneously produced efficiency gains that substantially exceeded either modality alone: correct diagnoses per minute increased by 40% and comment editing time fell by 67%, without compromising diagnostic accuracy. Notably, neither modality improved efficiency during guidance in US2, which makes the in-guidance efficiency gain under combined guidance in US3 the more striking result. Expert-distilled multimodal guidance can remove two distinct clinical workflow bottlenecks at once (visual search overhead and documentation burden) without compromising the diagnostic accuracy clinicians already achieve.

医学AI视觉引导多模态OCT诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。