arXiv:2512.19107cs.AI2025-12被引 2

用关键帧压缩提升手机界面操作意图识别效率,支持轻量化部署。

FC-MIR: A Mobile Screen Awareness Framework for Intent-Aware Recommendation based on Frame-Compressed Multimodal Trajectory Reasoning

  • 通过关键帧采样与自适应拼接减少视觉冗余,提升推理效率。
  • 在50%-60%压缩率下保持性能,支持端侧轻量部署。
  • 适用于移动智能助手、自动化任务代理等场景。

从移动端UI操作轨迹中识别用户意图,对推进UI理解与任务自动化代理至关重要。尽管多模态大语言模型(MLLMs)在视频理解任务中表现优异,但其在移动端实时部署受限于高计算开销和冗余帧处理效率低下。为此,我们提出FC-MIR框架:通过关键帧采样与自适应拼接,减少视觉冗余以提升推理效率,并集成先进闭源MLLM或微调模型(如Qwen3-VL)实现轨迹摘要与意图预测。我们进一步拓展任务范围,探索预测后的操作生成与搜索建议,并引入细粒度评估指标衡量摘要、预测与建议的实际效用。为严格评估,我们构建了涵盖UI-Agent(Agent-I)与真实用户交互(Person-I)场景的UI轨迹数据集。实验表明,该压缩方法在50%-60%压缩率下仍保持性能;闭源与微调MLLM均展现强意图摘要能力,支持潜在轻量化端侧部署。然而,MLLM在生成有用且“出人意料”的建议方面仍存不足,有待改进。最后,我们在真实场景中部署该框架,整合UI感知与UI-Agent代理,为该领域未来发展奠定基础。

原文摘要 · Abstract (English)

Identifying user intent from mobile UI operation trajectories is critical for advancing UI understanding and enabling task automation agents. While Multimodal Large Language Models (MLLMs) excel at video understanding tasks, their real-time mobile deployment is constrained by heavy computational costs and inefficient redundant frame processing. To address these issues, we propose the FC-MIR framework: leveraging keyframe sampling and adaptive concatenation, it cuts visual redundancy to boost inference efficiency, while integrating state-of-the-art closed-source MLLMs or fine-tuned models (e.g., Qwen3-VL) for trajectory summarization and intent prediction. We further expand task scope to explore generating post-prediction operations and search suggestions, and introduce a fine-grained metric to evaluate the practical utility of summaries, predictions, and suggestions. For rigorous assessment, we construct a UI trajectory dataset covering scenarios from UI-Agents (Agent-I) and real user interactions (Person-I). Experimental results show our compression method retains performance at 50%-60% compression rates; both closed-source and fine-tuned MLLMs demonstrate strong intent summarization, supporting potential lightweight on-device deployment. However, MLLMs still struggle with useful and "surprising" suggestions, leaving room for improvement. Finally, we deploy the framework in a real-world setting, integrating UI perception and UI-Agent proxies to lay a foundation for future progress in this field.

意图识别多模态模型轻量化部署移动智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。