让患者或医生通过多轮对话精准找到所需健康视频
Interactive Multi-Turn Retrieval for Health Videos

- 设计多轮对话式检索框架,支持逐步细化查询条件
- 在新构建的MHVRC数据集上超越现有基线模型
- 适合临床培训、康复指导等需要精细步骤匹配的场景
健康类教学视频日益普及,为临床培训、患者康复和健康教育提供了新可能,但现有检索系统多为单轮交互:用户提交一次查询,获得一个排序列表。这种模式在健康场景中脆弱,因信息需求常初始模糊,需通过体位、手部位置、禁忌症、设备或患者状况等后续约束才变得明确。本文提出面向健康视频的交互式多轮语义检索,并构建了MHVRC多轮健康视频检索语料库,结合VideoChat-Flash的视频描述与DeepSeek生成的查询优化结果。进一步提出DATR对话感知双阶段检索框架:先用类似CLIP的双编码器与稀疏帧采样进行高效粗检索,再通过多轮查询融合与轻量级交叉编码器对候选视频重新排序。在MHVRC上的实验显示,该方法持续优于强文本-视频检索基线;用户研究也表明,经过优化的多轮查询能更准确捕捉细粒度操作语义。本工作建立了基准并提供可扩展的技术方案。
原文摘要 · Abstract (English)
The growing availability of health-related instructional videos creates new opportunities for clinical training, patient rehabilitation, and health education, yet existing retrieval systems remain largely single-turn: a user submits one query and receives one ranked list. This interaction is brittle in health scenarios, where information needs are often vague at first and become clinically meaningful only after follow-up constraints such as posture, hand placement, contraindications, equipment, or patient condition are specified. We introduce interactive multi-turn semantic retrieval for health videos and construct MHVRC, a Multi-Turn Health Video Retrieval Corpus, by combining video-grounded descriptions from VideoChat-Flash with query refinements generated by DeepSeek. We further propose DATR, a Dialogue-Aware Two-Stage Retrieval framework. DATR first performs efficient coarse retrieval with a CLIP-style dual encoder and sparse frame sampling, then re-ranks the top candidates through multi-turn query fusion and a lightweight cross-encoder scoring module. Experiments on MHVRC show consistent gains over strong text-video retrieval baselines, while user studies indicate that refined multi-turn queries better capture fine-grained procedural semantics than single-turn annotations. The work establishes a benchmark and a scalable technical recipe for interactive health video retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。