提出高效对象级解释方法PhaseWin,计算量降为20%仍保持95%精度
PhaseWin Search Framework Enable Efficient Object-Level Interpretation
- 分阶段窗口搜索替代传统贪心选择,实现近线性复杂度
- 在20%计算预算下达成超过95%的解释忠实度
- 适用于目标检测与视觉定位任务,适合部署于实际场景
解释能力对理解对象级基础模型至关重要。近期基于子模集选择的方法虽具备高忠实度,但效率受限难以落地。为此,我们提出PhaseWin——一种新型相位窗口搜索算法,可在近线性复杂度下实现高忠实度区域归因。PhaseWin以分阶段粗到精搜索替代传统二次复杂度贪心选择,结合自适应剪枝、窗口化细粒度选择与动态监督机制,近似贪心行为同时大幅减少模型评估次数。理论上,在温和单调子模假设下,保留接近贪心的近似保证。实验表明,PhaseWin仅用20%计算预算即达到超过95%的贪心归因忠实度,并在目标检测与视觉定位任务中持续优于其他基线方法,使用Grounding DINO与Florence-2模型。PhaseWin建立了对象级多模态模型可扩展、高忠实度归因的新基准。
原文摘要 · Abstract (English)
Attribution is essential for interpreting object-level foundation models. Recent methods based on submodular subset selection have achieved high faithfulness, but their efficiency limitations hinder practical deployment in real-world scenarios. To address this, we propose PhaseWin, a novel phase-window search algorithm that enables faithful region attribution with near-linear complexity. PhaseWin replaces traditional quadratic-cost greedy selection with a phased coarse-to-fine search, combining adaptive pruning, windowed fine-grained selection, and dynamic supervision mechanisms to closely approximate greedy behavior while dramatically reducing model evaluations. Theoretically, PhaseWin retains near-greedy approximation guarantees under mild monotone submodular assumptions. Empirically, PhaseWin achieves over 95% of greedy attribution faithfulness using only 20% of the computational budget, and consistently outperforms other attribution baselines across object detection and visual grounding tasks with Grounding DINO and Florence-2. PhaseWin establishes a new state of the art in scalable, high-faithfulness attribution for object-level multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。