arXiv:2602.24142cs.CLcs.AI2026-02被引 5

提出CoME架构,让手机智能体更聪明地分步完成任务。

CoME: Empowering Channel-of-Mobile-Experts with Informative Hybrid-Capabilities Reasoning

论文配图:CoME: Empowering Channel-of-Mobile-Experts with Informative Hybrid-Capabilities Reasoning
图 1 · 摘自论文原文
  • 分四个专家模块,按任务阶段分别激活,实现能力解耦与协同
  • 在两个数据集上超越现有密集模型和MoE方法,准确率显著提升
  • 用信息增益引导推理过程,减少错误累积,适合复杂任务场景

移动智能体需自主执行用户指令,依赖屏幕摘要、子任务规划、动作决策与动作函数等混合能力。然而现有方法难以同时实现各能力的解耦增强与平衡整合。为此,我们提出通道式手机专家架构(CoME),包含四个对应不同推理阶段的独立专家,通过面向输出的激活机制,在每个阶段调用相应专家生成输出。为赋能混合能力推理,我们设计渐进式训练策略:Expert-FT 实现各专家能力的解耦增强;Router-FT 使专家激活与推理阶段对齐;CoT-FT 促进多能力间无缝协作与均衡优化。为缓解混合推理中的误差传播,提出基于信息增益的直接偏好优化(Info-DPO),利用信息增益评估每一步贡献,引导智能体生成更具信息量的推理路径。大量实验表明,CoME 在 AITZ 与 AMEX 数据集上均优于密集型移动智能体及 MoE 方法。

原文摘要 · Abstract (English)

Mobile Agents can autonomously execute user instructions, which requires hybrid-capabilities reasoning, including screen summary, subtask planning, action decision and action function. However, existing agents struggle to achieve both decoupled enhancement and balanced integration of these capabilities. To address these challenges, we propose Channel-of-Mobile-Experts (CoME), a novel agent architecture consisting of four distinct experts, each aligned with a specific reasoning stage, CoME activates the corresponding expert to generate output tokens in each reasoning stage via output-oriented activation. To empower CoME with hybrid-capabilities reasoning, we introduce a progressive training strategy: Expert-FT enables decoupling and enhancement of different experts' capability; Router-FT aligns expert activation with the different reasoning stage; CoT-FT facilitates seamless collaboration and balanced optimization across multiple capabilities. To mitigate error propagation in hybrid-capabilities reasoning, we propose InfoGain-Driven DPO (Info-DPO), which uses information gain to evaluate the contribution of each intermediate step, thereby guiding CoME toward more informative reasoning. Comprehensive experiments show that CoME outperforms dense mobile agents and MoE methods on both AITZ and AMEX datasets.

智能体移动推理混合能力专家系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。