提出跨视角角色不对称的手术视频预测模型,提升双镜协同建模能力。
CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction

- 设计角色不对称双流架构,按任务需求选择性传递跨视角信息。
- 在真实与模拟ERCP数据上,预测精度、运动一致性均优于现有方法。
- 适合需要多视角协作的医疗视觉建模研究者参考。
视觉世界模型通常仅从单一观察流学习未来动态,难以建模多个独立移动观察者的合作系统。本文研究母镜-子镜内镜逆行胰胆管造影(ERCP)中的挑战,其中两根柔性内镜提供互补但角色依赖的视角,且无校准的立体关系。不同于传统对称多视角融合,本文提出角色不对称双视角未来预测,根据预测目标及其空间需求选择性传递跨视角证据。我们提出CrossScope,一种双流手术世界模型,在保留各视角特异性专家的同时,通过几何引导的残差交互实现目标特定证据路由。模型学习两种互补通信方向:母镜的几何运动线索指导子镜未来动态,而姿态对齐的子镜外观仅在建立有效空间对应时支持母镜预测。该设计使每个镜头仅贡献任务相关证据,同时不损害其视角特异性表征。为评估该问题,我们建立配对双镜基准,包含同步的模拟与真实ERCP片段,评估涵盖视觉保真度、结构保持、目标定位与运动一致性。实验表明,CrossScope持续优于强基线,验证了角色感知证据路由在多观察者视觉建模中的重要性。
原文摘要 · Abstract (English)
Visual world models typically learn future dynamics from a single observation stream, limiting their ability to model cooperative systems with multiple independently moving observers. We investigate this challenge in Mother--Child endoscopic retrograde cholangiopancreatography (ERCP), where two flexible scopes provide complementary yet role-dependent views without a calibrated stereo relationship. Unlike conventional multi-view fusion that assumes symmetric information exchange, we formulate \textbf{role-asymmetric dual-scope future prediction}, where cross-view evidence is selectively transferred according to the prediction target and its underlying spatial requirements. We propose \textbf{CrossScope}, a dual-stream surgical world model that preserves view-specific experts while enabling target-specific evidence routing through geometry-guided residual interactions. CrossScope learns two complementary communication directions: geometric motion cues from the Mother view guide Child-view future dynamics, while pose-aligned Child appearance supports Mother-view prediction only when valid spatial correspondence is established. This design allows each scope to contribute task-relevant evidence without compromising its view-specific representation. To evaluate this problem, we establish a paired dual-scope benchmark comprising synchronized phantom and real-world ERCP episodes, with evaluations assessing visual fidelity, structural preservation, target localization, and motion consistency. Experiments demonstrate that CrossScope consistently outperforms strong surgical video generation baselines, validating the importance of role-aware evidence routing for multi-observer visual world modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。