用少参数多任务模型,让胶囊内镜更省电、更智能
Seeing More with Less: Video Capsule Endoscopy with Multi-Task Learning
- 一个模型同时完成定位与病变检测,减少重复计算
- 仅100万参数却达93.63%定位准确率和87.48%检测准确率
- 适合资源受限的医疗边缘设备,推动智能内镜发展
视频胶囊内镜在胃肠道小肠检查中日益重要,但其紧凑传感器设备的电池寿命短仍是难题。引入人工智能可实现智能实时决策,降低能耗、延长续航,但受限于数据稀疏和设备算力,模型规模难以扩大。本文提出一种多任务神经网络,在单一模型中集成小肠精准定位与病灶检测功能,全程控制参数总量以确保部署可行性。基于新发布的Galar数据集,首次实现多任务联合训练,并结合Viterbi解码进行时序分析。结果优于现有单任务模型,定位准确率达93.63%,异常检测准确率达87.48%,模型仅需100万参数即超越当前基线。
原文摘要 · Abstract (English)
Video capsule endoscopy has become increasingly important for investigating the small intestine within the gastrointestinal tract. However, a persistent challenge remains the short battery lifetime of such compact sensor edge devices. Integrating artificial intelligence can help overcome this limitation by enabling intelligent real-time decision-making, thereby reducing the energy consumption and prolonging the battery life. However, this remains challenging due to data sparsity and the limited resources of the device restricting the overall model size. In this work, we introduce a multi-task neural network that combines the functionalities of precise self-localization within the gastrointestinal tract with the ability to detect anomalies in the small intestine within a single model. Throughout the development process, we consistently restricted the total number of parameters to ensure the feasibility to deploy such model in a small capsule. We report the first multi-task results using the recently published Galar dataset, integrating established multi-task methods and Viterbi decoding for subsequent time-series analysis. This outperforms current single-task models and represents a significant advance in AI-based approaches in this field. Our model achieves an accuracy of 93.63% on the localization task and an accuracy of 87.48% on the anomaly detection task. The approach requires only 1 million parameters while surpassing the current baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。