提出分层提示框架,让模型学新器械时不忘旧技能,还能提升老技能。
Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt Framework
- 用分层提示树复用历史知识,简化新器械学习
- 通过自反思机制优化旧知识,避免遗忘且提升精度
- 适配CNN和Transformer模型,在两个公开数据集上显著领先
为持续提升外科视频场景解析中模型的适应性,现有研究通过增量学习逐步掌握更多手术器械的分割。然而,以往工作忽视了正向知识迁移(过去知识帮助学习新类)和反向知识迁移(学习新类改善旧知识)的潜力。本文提出一种自反思分层提示框架,旨在同时实现高效学习新器械、强化已有器械识别能力,并避免旧知识遗忘。该框架基于冻结的预训练模型,动态添加器械感知提示。为促进正向迁移,将提示组织成层次化提示解析树,以共享提示为根节点,部分共享提示为中间节点,专属提示为叶节点,暴露可复用的历史知识;为激发反向迁移,采用有向加权图传播进行自反思优化,分析树中记录的知识关联,增强表征能力而不引发灾难性遗忘。该框架适用于基于CNN和先进Transformer的基座模型,在两个公开基准上分别取得超过5%和11%的性能提升。
原文摘要 · Abstract (English)
To continuously enhance model adaptability in surgical video scene parsing, recent studies incrementally update it to progressively learn to segment an increasing number of surgical instruments over time. However, prior works constantly overlooked the potential of positive forward knowledge transfer, i.e., how past knowledge could help learn new classes, and positive backward knowledge transfer, i.e., how learning new classes could help refine past knowledge. In this paper, we propose a self-reflection hierarchical prompt framework that unlocks the power of positive forward and backward knowledge transfer in class incremental segmentation, aiming to proficiently learn new instruments, improve existing skills of regular instruments, and avoid catastrophic forgetting of old instruments. Our framework is built on a frozen, pre-trained model that adaptively appends instrument-aware prompts for new classes throughout training episodes. To enable positive forward knowledge transfer, we organize instrument prompts into a hierarchical prompt parsing tree with the instrument-shared prompt partition as the root node, n-part-shared prompt partitions as intermediate nodes and instrument-distinct prompt partitions as leaf nodes, to expose the reusable historical knowledge for new classes to simplify their learning. Conversely, to encourage positive backward knowledge transfer, we conduct self-reflection refining on existing knowledge by directed-weighted graph propagation, examining the knowledge associations recorded in the tree to improve its representativeness without causing catastrophic forgetting. Our framework is applicable to both CNN-based models and advanced transformer-based foundation models, yielding more than 5% and 11% improvements over the competing methods on two public benchmarks respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。