用自监督Transformer解决微创手术中光纤传感器的故障容错问题
Self-Supervised Mask-Aware Transformers for Fault-Tolerant FBG Force Sensing in Minimally Invasive Surgical Robotics

- 设计掩码感知Transformer,动态建模通道可用性
- 8通道下故障时误差仅升至0.0126N,优于255个模型组合
- 单次推理即可输出置信度,适合实时医疗场景
在微创手术机器人中,导管级光纤布喇格光栅(FBG)传感器因可多通道复用实现多维力估计而具有潜力。但其部署面临两大挑战:复杂形变下的固有非线性交叉轴耦合,以及受限空间内光纤断裂导致的间歇性通道丢失,二者叠加严重降低力估计精度。现有容错方法依赖组合式模型库,随通道数呈指数增长,且需昂贵的每模式校准。本文提出统一的自监督掩码感知Transformer,显式建模通道可用性,实现多样动态传感器故障下的渐进式退化。编码器通过无标签数据流上的掩码通道重建预训练,再以均衡的干净与损坏视图目标及动态扰动课程进行微调。此外,通过异方差高斯负对数似然训练的并行不确定性头,在单次前向传播中预测各轴置信度,避免多轮集成开销。在8通道导管级FBG数据集上,该单一模型在正常情况下达到0.0066N的均方根误差,4通道严重失效时仅升至0.0126N,显著优于包含255个特定模式神经网络的模型库(0.0154N),同时免除模式专属校准。
原文摘要 · Abstract (English)
In minimally invasive surgical robotics, catheter-scale Fiber Bragg Grating (FBG) sensors are promising due to their ability to estimate multi-dimensional forces by multiplexing several optical channels. However, deploying these compact multi-channel sensors introduces two critical engineering challenges: inherent nonlinear cross-axis coupling during complex deformations, and intermittent channel dropouts caused by fiber fractures in constrained workspaces. These compounding issues severely degrade force estimation. Existing fault-tolerant approaches rely on combinatorial model banks, which scale exponentially with the channel count and demand prohibitively expensive per-pattern calibration. In this paper, we propose a unified, self-supervised mask-aware Transformer that explicitly models channel availability to enable graceful degradation under diverse and dynamic sensor failures. The encoder is pretrained via masked-channel reconstruction on unlabeled data streams and fine-tuned for force regression using a balanced clean-and-corrupted-view objective alongside a dynamic corruption curriculum. Furthermore, a parallel uncertainty head, trained via heteroscedastic Gaussian negative log-likelihood, predicts per-axis confidence in a single forward pass, circumventing the overhead of multi-pass ensembles. Evaluated on a catheter-scale 8-channel FBG dataset, our single unified model achieves a nominal Root Mean Square Error (RMSE) of 0.0066~N and degrades gracefully to 0.0126~N under severe 4-channel failures. This significantly outperforms a comprehensive model bank of 255 per-pattern neural networks (0.0154~N at 4-channel loss) while eliminating pattern-specific calibration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。