arXiv:2601.11631cs.CV2026-01被引 2

通过坐标压缩提升多轮GUI智能体的决策效率与精度

Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI Agents

  • 基于多轮交互构建坐标感知的视觉压缩机制,聚焦关键区域
  • 在4个基准上实现最高55%的令牌压缩率和3.8倍训练加速
  • 适合需要高效长程上下文处理的GUI自动化任务研究者

多轮GUI智能体通过序列决策完成复杂任务,但交互历史积累导致上下文膨胀严重。现有方法或通过截断牺牲长期上下文,或通过标记剪枝破坏空间结构。本文提出坐标压缩策略优化(CCPO),将视觉压缩与策略优化结合。CCPO引入坐标感知空间压缩(CASC),聚合多轮回放中的坐标信息,识别目标相关区域,并逐步缩小历史注意力范围至关键视觉区域。CASC从多轮交互中自适应构建注意力边界,集中计算于场景中最相关信息区域。此外,设计基于距离的优势函数,提供细粒度学习信号,而非仅依赖二值正确性,从而提升定位准确率与压缩质量。大量实验表明,CCPO在四个基准上达到当前最优性能,实现最高55%的令牌压缩率和3.8倍训练速度提升。

原文摘要 · Abstract (English)

Multi-turn GUI agents enable complex task completion through sequential decision-making, but suffer from severe context inflation as interaction history accumulates. Existing strategies either sacrifice long-term context via truncation or compromise spatial structure through token pruning. In this paper, we propose Coordinate Compression Policy Optimization (CCPO), an efficient policy optimization framework that couples visual compression with policy optimization for multi-turn GUI agents. CCPO introduces Coordinate-Aware Spatial Compression (CASC), which aggregates coordinates from multiple rollouts to capture target-relevant regions and progressively narrow historical attention around key visual areas. From interactions across rollouts, CASC adaptively constructs attention boundaries that concentrate computation on the most informative regions of the scene. We further design a Distance-Based Advantage that provides fine-grained learning signals based on distance rather than binary correctness, improving both grounding accuracy and compression quality. Extensive experiments demonstrate that CCPO achieves SOTA performance across four benchmarks with up to 55% token compression and 3.8$\times$ training speedup.

GUI智能体视觉压缩策略优化多轮决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。