用物理与语义引导的强化学习,让扩散模型更智能地清除手术烟雾。
PhySe-RPO: Physics and Semantics Guided Relative Policy Optimization for Diffusion-Based Surgical Smoke Removal
- 将去烟过程变为随机策略,支持多路径探索与无需评判器的优化
- 在真实机器人手术数据上实现物理一致、语义准确的去烟效果
- 适合缺乏成对标注的医疗视觉任务,尤其适用于手术视频增强
手术烟雾严重降低术中视频质量,遮挡解剖结构,影响手术判断。现有基于学习的去烟方法依赖稀缺的成对监督信号,且采用确定性恢复流程,难以在真实手术条件下进行探索或强化驱动优化。本文提出 PhySe-RPO,一种通过物理与语义引导的相对策略优化进行优化的扩散恢复框架。核心思想是将确定性恢复转化为随机策略,通过群体相对优化实现轨迹级探索和无评判器更新。物理引导奖励确保光照与色彩一致性,基于 CLIP 的视觉概念语义奖励促进无烟且解剖结构一致的恢复。结合无参考感知约束,PhySe-RPO 在合成与真实机器人手术数据集上均生成物理一致、语义忠实且临床可解释的结果,为有限成对监督下的鲁棒扩散恢复提供了系统性解决方案。
原文摘要 · Abstract (English)
Surgical smoke severely degrades intraoperative video quality, obscuring anatomical structures and limiting surgical perception. Existing learning-based desmoking approaches rely on scarce paired supervision and deterministic restoration pipelines, making it difficult to perform exploration or reinforcement-driven refinement under real surgical conditions. We propose PhySe-RPO, a diffusion restoration framework optimized through Physics- and Semantics-Guided Relative Policy Optimization. The core idea is to transform deterministic restoration into a stochastic policy, enabling trajectory-level exploration and critic-free updates via group-relative optimization. A physics-guided reward imposes illumination and color consistency, while a visual-concept semantic reward learned from CLIP-based surgical concepts promotes smoke-free and anatomically coherent restoration. Together with a reference-free perceptual constraint, PhySe-RPO produces results that are physically consistent, semantically faithful, and clinically interpretable across synthetic and real robotic surgical datasets, providing a principled route to robust diffusion-based restoration under limited paired supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。