评测医疗视频AI何时该回答,而非只看对错。
MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding

- 构建时间感知的医疗视频评测基准,支持流式和主动预警场景。
- 涵盖5419个问答实例,覆盖回顾、当前、未来和主动四种时间设定。
- 揭示主流模型在实时决策中表现显著下降,适合临床部署研究者使用。
现有医疗视频基准主要评估模型答案是否正确,却很少检验回答时机是否恰当。在真实临床中,AI系统不仅要预测内容,还需判断何时回应、延迟判断或主动预警。这导致评测与实际部署间存在关键差距。我们提出MedStreamBench,一个面向时间感知医疗视频理解的基准。该基准整合22个医疗数据集,包含5,419个问答实例,覆盖四种时间场景:回顾、当前、未来和主动。不同于传统假设全视频可访问的基准,MedStreamBench限制模型仅能获取时序受限的证据窗口,并支持单轮与流式评估。我们进一步引入主动监控场景,要求模型判断何时触发临床警报。除答案正确性外,还评估响应速度与证据后稳定性。在主流通用及医疗视觉语言模型上的实验显示,离线识别与时空决策之间存在显著差距,模型在流式和主动场景中性能大幅下降。基准已开源:https://huggingface.co/datasets/Venn2024/MedStreamBench。
原文摘要 · Abstract (English)
Existing medical video benchmarks primarily evaluate whether a model produces the correct answer, but rarely assess whether it answers at the right time. In real clinical settings, AI systems must decide not only what to predict, but also when to answer, defer judgment, or proactively raise alerts. This creates a critical gap between benchmark evaluation and deployment requirements. We present MedStreamBench, a benchmark for time-aware medical video understanding. MedStreamBench integrates 22 medical datasets and 5,419 QA instances across four temporal settings: retrospective, present, future, and proactive. Unlike conventional benchmarks that assume full-video access, MedStreamBench restricts models to temporally bounded evidence windows and supports both single-turn and streaming evaluation. We further introduce a proactive monitoring setting that requires models to determine whether and when clinically relevant alerts should be triggered. Beyond answer correctness, MedStreamBench evaluates temporal behavior through responsiveness and post-evidence stability. Experiments on leading general-purpose and medical vision-language models reveal a substantial gap between offline recognition and temporally grounded decision-making, with performance dropping markedly in streaming and proactive settings. Our benchmark is available at https://huggingface.co/datasets/Venn2024/MedStreamBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。