Mamba模型处理文档二值化时可避免弱笔画信息丢失
DeepMine-Mamba: Mitigating Information Dilution in Mamba-Based State Space Models for Document Image Binarization

- 引入抗稀释门机制,动态恢复易丢失的细弱笔画特征
- 在DIBCO/H-DIBCO上实现高平均F-measure与准确率
- 适合处理模糊、断裂、低对比度文字的场景
文档图像二值化旨在分离前景文本与退化背景,同时保留细小、断裂及低对比度笔画。尽管深度学习已提升性能,现有方法多依赖卷积、Transformer或生成架构,基于Mamba的状态空间模型尚未被充分探索。本文研究Mamba特征传播机制,发现其长程建模可能稀释微弱前景信号,如淡墨痕迹、碎片化字符与边界敏感笔画。为此提出DeepMine-Mamba框架,设计新型抗稀释门,估计传播过程中的特征变化,选择性恢复笔画敏感局部响应,抑制冗余背景增强。在严格留一年外测试协议下,于DIBCO/H-DIBCO基准上取得竞争性整体表现,跨年平均F-measure与精确率均较高。消融实验证明抗稀释门是缓解传播导致前景稀释、提升笔画保真度的关键组件。
原文摘要 · Abstract (English)
Document image binarization aims to separate foreground text from degraded backgrounds while preserving thin, broken, and low-contrast strokes. Although deep learning methods have improved binarization performance, most existing approaches rely on convolutional, transformer-based, or generative architectures, while Mamba-based state space models remain largely unexplored for this task. In this work, we investigate Mamba-based feature propagation and observe that direct state-space propagation may dilute weak foreground cues during long-range modeling, especially faint ink traces, fragmented characters, and boundary-sensitive stroke details. To address this problem, we propose DeepMine-Mamba, a Mamba-based binarization framework equipped with a novel Anti-Dilution Gate that estimates propagation-induced feature changes and selectively restores stroke-sensitive local responses while suppressing unnecessary background enhancement. Experiments on DIBCO/H-DIBCO benchmarks under a strict leave-one-year-out protocol show that DeepMine-Mamba achieves competitive overall performance, with strong average FM and Fps across benchmark years. Ablation results further show that the Anti-Dilution Gate is the key component for mitigating propagation-induced foreground dilution and improving stroke preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。