arXiv:2607.17656cs.CV2026-07

解决视觉变压器中边缘小区域被忽略的问题,提升关键信息捕捉能力。

BMFA: Boundary-Minority Free-Energy Adaptive Screening

论文配图:BMFA: Boundary-Minority Free-Energy Adaptive Screening
图 1 · 摘自论文原文
  • 通过自由能分析发现小区域虽弱但贡献大,传统平均法会遗漏
  • 提出BMFA方法,在5.79%叶节点下将误差降低至0.261,图像边缘误差降为0.526
  • 适合关注精度与细节的视觉模型优化,尤其对小目标检测有帮助

视觉变压器仅在粗粒度标记摘要保留指数注意力聚合所需证据时才能高效处理空间冗余标记。我们识别出一种边界少数低估失败现象:空间上微小但响应高的区域虽主导吉布斯质量,却几乎无法被块均值察觉。通过归一化对数-均-指数自由能与均值摘要之间的差异形式化该失败,证明即使空间支持和均值贡献趋近于零,少数吉布斯质量仍可不衰减,并刻画了有限阶矩修正的局限性。基于此分析,提出边界少数自由能自适应筛选(BMFA),构建分层分段常数逼近,递归根据局部自由能下界增量精修块。控制合成测试、COCO与LVIS诊断探针、闭环DeiT-Tiny评估及ImageNet-1K实验建立一致证据链。BMFA将平均合成低估从2.582降至0.261(叶节点比5.794%),将COCO图像边缘均值差距从2.254降至0.526,且在55.861%叶节点比下保持71.520% ImageNet Top-1准确率。当前原型在完成全QK计算后评估选择质量,故报告叶节点比表征表示粒度而非验证的稀疏内核加速效果。

原文摘要 · Abstract (English)

Vision Transformers process spatially redundant tokens efficiently only when coarse token summaries preserve the evidence required by exponential attention aggregation. We identify a boundary-minority underestimation failure in which a spatially small, high-response region contributes dominant Gibbs mass while remaining nearly invisible to a block mean. We formalize the failure through the discrepancy between normalized log-mean-exp free energy and mean summarization, prove that minority Gibbs mass can remain non-vanishing as its spatial support and mean contribution vanish, and characterize the limitations of finite-order moment corrections. Building on the resulting analysis, we introduce Boundary-Minority Free-Energy Adaptive Screening (BMFA), which constructs a hierarchical piecewise-constant approximation and recursively refines blocks according to a computable lower-bound increment of local free energy. Controlled synthetic tests, COCO and LVIS diagnostic probes, closed-loop DeiT-Tiny evaluations, and ImageNet-1K experiments establish a consistent evidence chain. BMFA reduces the mean synthetic underestimate from 2.582 to 0.261 at a 5.794% leaf ratio, lowers the COCO image-edge mean gap from 2.254 to 0.526, and preserves 71.520% ImageNet Top-1 accuracy at a 55.861% leaf ratio. The current prototype evaluates selection quality after full QK computation; the reported leaf ratio therefore characterizes representation granularity rather than verified sparse-kernel speedup.

视觉变压器自由能稀疏注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。