arXiv:2604.10916cs.CVcs.AI2026-04

构建首个面向超声操作流程的视频问答基准,助力智能超声系统发展

ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding

论文配图:ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
图 1 · 摘自论文原文
  • 设计视频QA数据集,聚焦操作目标推理、伪影处理与流程规划三类能力
  • 514段视频+514题测试,发现大模型在故障诊断上仍远落后于文本基线
  • 适合医学AI、机器人辅助超声及临床训练系统研究者使用

超声采集依赖熟练的探头操作与实时调整。视觉语言模型(VLMs)有望实现自主超声系统,但现有基准仅评估静态图像,缺乏对动态流程理解的评测。我们提出ReXSonoVQA,一个包含514个视频片段和514个问题(249道多选题,265道开放题)的视频问答基准,针对三大能力:动作-目标推理、伪影识别与优化、流程上下文与规划。对Gemini 3 Pro、Qwen3.5-397B、LLaVA-Video-72B和Seed 2.0 Pro的零样本评估显示,尽管模型可提取部分流程信息,但在故障排查类问题上表现不佳,改进微弱,暴露出因果推理能力的局限性。ReXSonoVQA为超声教学、引导与机器人自动化系统的感知能力提升提供评测基础。

原文摘要 · Abstract (English)

Ultrasound acquisition requires skilled probe manipulation and real-time adjustments. Vision-language models (VLMs) could enable autonomous ultrasound systems, but existing benchmarks evaluate only static images, not dynamic procedural understanding. We introduce ReXSonoVQA, a video QA benchmark with 514 video clips and 514 questions (249 MCQ, 265 free-response) targeting three competencies: Action-Goal Reasoning, Artifact Resolution & Optimization, and Procedure Context & Planning. Zero-shot evaluation of Gemini 3 Pro, Qwen3.5-397B, LLaVA-Video-72B, and Seed 2.0 Pro shows VLMs can extract some procedural information, but troubleshooting questions remain challenging with minimal gains over text-only baselines, exposing limitations in causal reasoning. ReXSonoVQA enables developing perception systems for ultrasound training, guidance, and robotic automation.

视频问答超声影像多模态医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。