构建真实医疗长视频理解基准,挑战模型在冗余中找关键证据的能力。
MedHorizon: Towards Long-context Medical Video Understanding in the Wild

- 提出新基准MedHorizon,保留759小时完整手术视频。
- 平均仅0.166%帧为有效证据,模型需从噪声中检索并推理。
- 揭示当前模型在长程推理与注意力聚焦上的根本短板。
医学多模态大模型在图像和短视频理解上已取得进展,但临床实际需要对完整手术过程进行理解。与通用长视频不同,医疗操作包含高度冗余的解剖视角,而关键证据具有时间稀疏、空间细微、上下文依赖的特点。现有基准常假设证据已被图像、短片段或预分割视频定位,导致‘先检索后推理’问题未被充分测试。我们提出MedHorizon——一个真实场景下的长时序医疗视频理解基准。该数据集保留了759小时的完整临床手术视频,提供1,253个基于证据的多选题,联合评估稀疏证据理解与多跳临床推理能力。其证据极为稀疏,平均仅有0.166%的帧为有效证据,要求模型在嘈杂的手术流中搜索、解释并聚合信息。我们评估了代表性通用领域、医疗领域及长视频多模态大模型,最佳模型准确率仅为41.1%,表明当前系统距离稳健的全流程理解仍有巨大差距。进一步分析发现:性能不随帧数增加而可靠提升;证据检索与临床解读是主要瓶颈;根源在于弱程序推理能力与冗余下的注意力漂移;通用采样方法仅部分平衡局部细节与全局覆盖。MedHorizon为测试模型在完整临床流程中检索稀疏证据并推理提供了严格基准。
原文摘要 · Abstract (English)
Medical multimodal large language models (MLLMs) have advanced image understanding and short-video analysis, but real clinical review often requires full-procedure video understanding. Unlike general long videos, medical procedures contain highly redundant anatomical views, while decisive evidence is temporally sparse, spatially subtle, and context dependent. Existing benchmarks often assume this evidence has already been localized through images, short clips, or pre-segmented videos, leaving the retrieval-before-reasoning problem under-tested. We introduce MedHorizon, an in-the-wild benchmark for long-context medical video understanding. MedHorizon preserves 759 hours of full-length clinical procedures and provides 1,253 evidence-grounded multiple-choice questionsthat jointly evaluate sparse evidence understanding and multi-hop clinical reasoning. Its evidence is extremely sparse, with only 0.166% evidence frames on average, requiring models to search noisy procedural streams before interpreting and aggregating findings. We evaluate representative general-domain, medical-domain, and long-video MLLMs. The best model reaches only 41.1% accuracy, showing that current systems remain far from robust full-procedure understanding. Further analysis yields four key findings: performance does not scale reliably with more frames, evidence retrieval and clinical interpretation remain primary bottlenecks; these bottlenecks are rooted in weak procedural reasoning and attention drift under redundancy, and generic sampling methods only partially balances local detail with global coverage. MedHorizon provides a rigorous testbed for MLLMs that retrieve sparse evidence and reason over complete clinical workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。