用数字孪生+强化学习训练大模型,提升手术视频问答的多步推理能力
Training LLMs with Reinforcement Learning over Digital Twin Representations for Reasoning-Intensive Surgical VideoQA

- 通过数字孪生表示解耦视觉感知与推理过程
- 在三个复杂度等级的2000对问答上达到最优性能
- 适合需要精准医疗推理的临床研究与AI辅助决策场景
手术视频问答需在语义、空间和时间维度进行多步推理。现有方法将视频压缩为离散标记表示,并将视觉感知与推理耦合,破坏了连续时空关系,限制了多步推理能力。本文提出一种基于强化学习的框架,让大语言模型在由手术基础模型构建的数字孪生表示上操作,实现感知与推理的解耦。同时引入帧、时间窗口和手术阶段的分层表示,并包含概率不确定性估计。提出一种新奖励机制,结合格式验证、临床合理性评估和不确定性校准进行训练。为验证该方法,构建了包含2000个问答对的REAL-Colon-Reason基准,覆盖三个复杂度层级。在REAL-Colon-Reason及两个现有手术视频问答基准(REAL-Colon-VQA、EndoVis18-VQA)上均达到当前最优表现。
原文摘要 · Abstract (English)
Surgical video question answering requires multi-step reasoning across semantic, spatial, and temporal dimensions. Existing methods architecturally compress videos into discrete token representations and couple visual perception with reasoning. This approach fragments continuous spatial-temporal relationships and has been shown to restrict multi-step reasoning capabilities. We introduce a reinforcement learning (RL) framework that trains large language models (LLMs) to decouple perception from reasoning by operating over digital twin representations constructed from surgical foundation models. Additionally, we introduce hierarchical representations across frame, temporal window, and procedure levels with probabilistic uncertainty estimates. Finally, we propose a novel reward that combines format validation with accuracy assessment through clinical plausibility evaluation and uncertainty-aware calibration for training. To demonstrate the capabilities of this approach, we introduce REAL-Colon-Reason, a colonoscopic benchmark with 2000 question-answer pairs across three complexity levels. We achieve state-of-the-art performance on REAL-Colon-Reason and two existing surgical VideoQA benchmarks REAL-Colon-VQA and EndoVis18-VQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。