arXiv:2504.11406cs.CVcs.AI2025-04被引 1

用专家标记初始化多层元胞自动机,实现低资源下的高效显著性检测。

Multi-level Cellular Automata for FLIM networks

  • 用专家绘制的标记直接学习卷积编码器,无需反向传播。
  • 在两个医学数据集上达到主流深度模型水平,参数量远低于轻量级模型。
  • 适合计算资源有限的医疗场景,尤其适用于缺乏标注数据的地区。

深度学习显著性物体检测(deep SOD)及更广泛的深度学习领域面临标注数据不足与网络结构复杂的问题,这一挑战在发展中国家的医疗应用中尤为突出。结合现代与经典方法可兼顾性能与实用性。特征从图像标记中学习(FLIM)方法使专家通过手绘标记设计卷积编码器,滤波器直接由标注学习得到。近期研究表明,将FLIM编码器与自适应解码器结合可构建轻量级网络,参数量远小于轻量模型且无需反向传播。元胞自动机(CA)在数据稀缺场景表现优异,但需恰当初始化——通常依赖用户输入、先验或随机性。本文提出两者的实用融合:利用FLIM网络以专家知识初始化CA状态,无需逐图交互。通过解码FLIM网络各层级特征,可同时初始化多个CA,构建多层级框架。该方法整合不同网络层级的层次化知识,将多个显著图融合为高质量最终输出,形成CA集成。在两个具有挑战性的医学数据集上的基准测试表明,该多层CA方法在性能上媲美现有主流深度模型。

原文摘要 · Abstract (English)

The necessity of abundant annotated data and complex network architectures presents a significant challenge in deep-learning Salient Object Detection (deep SOD) and across the broader deep-learning landscape. This challenge is particularly acute in medical applications in developing countries with limited computational resources. Combining modern and classical techniques offers a path to maintaining competitive performance while enabling practical applications. Feature Learning from Image Markers (FLIM) methodology empowers experts to design convolutional encoders through user-drawn markers, with filters learned directly from these annotations. Recent findings demonstrate that coupling a FLIM encoder with an adaptive decoder creates a flyweight network suitable for SOD, requiring significantly fewer parameters than lightweight models and eliminating the need for backpropagation. Cellular Automata (CA) methods have proven successful in data-scarce scenarios but require proper initialization -- typically through user input, priors, or randomness. We propose a practical intersection of these approaches: using FLIM networks to initialize CA states with expert knowledge without requiring user interaction for each image. By decoding features from each level of a FLIM network, we can initialize multiple CAs simultaneously, creating a multi-level framework. Our method leverages the hierarchical knowledge encoded across different network layers, merging multiple saliency maps into a high-quality final output that functions as a CA ensemble. Benchmarks across two challenging medical datasets demonstrate the competitiveness of our multi-level CA approach compared to established models in the deep SOD literature.

显著性检测医学图像轻量模型元胞自动机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。