arXiv:2508.04801cs.CV2025-08被引 4

构建多语言耳鼻喉内镜分析挑战赛,支持细粒度分类与跨模态检索

ACM Multimedia Grand Challenge on ENT Endoscopy Analysis

  • 构建双语标注数据集,涵盖解剖区域与病变状态标签
  • 设计三类基准任务,实现图像/文本到图像的跨模态检索
  • 面向临床需求,适合医学AI与多模态检索研究者使用

自动化分析耳鼻喉内镜图像对临床诊疗至关重要,但受限于设备与操作者差异、病灶细微且局部化,以及侧别、声带状态等精细区分。除分类外,临床还需可靠地检索相似病例,包括视觉和简明文本描述。现有公开基准难以支持这些能力。为此,我们提出ENTRep——2025年ACM多媒体大会耳鼻喉内镜分析挑战赛,整合细粒度解剖分类与跨模态图像-图像、文本-图像检索,支持中英双语临床标注。数据集包含专家标注图像,标记解剖部位及正常/异常状态,并附双语叙事描述。定义三项基准任务,标准化提交流程,采用服务器端评分在公有与私有测试集上评估性能。报告优胜团队结果并提供深入讨论。

原文摘要 · Abstract (English)

Automated analysis of endoscopic imagery is a critical yet underdeveloped component of ENT (ear, nose, and throat) care, hindered by variability in devices and operators, subtle and localized findings, and fine-grained distinctions such as laterality and vocal-fold state. In addition to classification, clinicians require reliable retrieval of similar cases, both visually and through concise textual descriptions. These capabilities are rarely supported by existing public benchmarks. To this end, we introduce ENTRep, the ACM Multimedia 2025 Grand Challenge on ENT endoscopy analysis, which integrates fine-grained anatomical classification with image-to-image and text-to-image retrieval under bilingual (Vietnamese and English) clinical supervision. Specifically, the dataset comprises expert-annotated images, labeled for anatomical region and normal or abnormal status, and accompanied by dual-language narrative descriptions. In addition, we define three benchmark tasks, standardize the submission protocol, and evaluate performance on public and private test splits using server-side scoring. Moreover, we report results from the top-performing teams and provide an insight discussion.

医学影像多模态检索细粒度分类双语标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。