arXiv:2508.19679cs.AI2025-08ACL被引 6

让手机智能体学会在关键时主动求助,提升安全性和成功率。

InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning

  • 用强化学习训练模型,在行动前主动判断是否需人类确认。
  • 在新基准上查询成功率达46.8%提升,整体表现最佳。
  • 适合需要高安全性的人机协作场景,如智能助手、自动驾驶。

视觉语言模型(VLM)的进步使移动智能体能够基于人类指令感知并交互真实环境。然而,当前完全自主模式在模型理解或推理能力不足时存在潜在安全风险。为此,我们首先提出 extbf{InquireBench},一个全面评估移动智能体在安全交互与主动询问用户方面能力的基准,涵盖5类22子类,现有大多数VLM基智能体表现接近零分。本文旨在构建一种能主动在关键决策点寻求人类确认的交互系统。为此,我们提出 extbf{InquireMobile},一种受强化学习启发的新模型,包含两阶段训练策略和交互式预动作推理机制。最终,该模型在InquireBench上实现查询成功率46.8%的提升,并取得现有基线中最佳的整体成功率。所有数据集、模型与评估代码将开源,以推动学术界与产业界发展。

原文摘要 · Abstract (English)

Recent advances in Vision-Language Models (VLMs) have enabled mobile agents to perceive and interact with real-world mobile environments based on human instructions. However, the current fully autonomous paradigm poses potential safety risks when model understanding or reasoning capabilities are insufficient. To address this challenge, we first introduce \textbf{InquireBench}, a comprehensive benchmark specifically designed to evaluate mobile agents' capabilities in safe interaction and proactive inquiry with users, encompassing 5 categories and 22 sub-categories, where most existing VLM-based agents demonstrate near-zero performance. In this paper, we aim to develop an interactive system that actively seeks human confirmation at critical decision points. To achieve this, we propose \textbf{InquireMobile}, a novel model inspired by reinforcement learning, featuring a two-stage training strategy and an interactive pre-action reasoning mechanism. Finally, our model achieves an 46.8% improvement in inquiry success rate and the best overall success rate among existing baselines on InquireBench. We will open-source all datasets, models, and evaluation codes to facilitate development in both academia and industry.

人机协作强化学习智能体安全交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。