提出混合检测框架,提升移动端智能体操作安全性
OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
- 结合形式化验证与视觉语言模型,双重检测系统漏洞与上下文风险
- 在真实操作轨迹上实现10%-30%性能提升,显著增强安全检测能力
- 适用于开发更安全的自动化移动应用代理,研究者与开发者皆可参考
由视觉-语言模型驱动的计算机操作智能体在移动平台等数字环境中展现出类人能力。然而,其可能引发系统入侵和隐私泄露等安全隐患,成为数字自动化发展的重要障碍。现有方法难以有效覆盖移动环境复杂多变的操作空间,安全检测仍严重不足。为此,我们构建了MobileRisk-Live——一个动态沙盒环境及其配套的安全检测基准,包含真实操作轨迹与细粒度标注。基于此,提出OS-Sentinel:一种融合形式化验证器(检测显式系统违规)与视觉语言模型上下文判断器(评估行为上下文风险)的新型混合安全检测框架。实验表明,该框架在多个指标上相较现有方法提升10%-30%。深入分析揭示了提升安全性的关键路径,为构建更可靠、更安全的自主移动代理提供重要启示。代码与数据已公开于https://qiushisun.github.io/OS-Sentinel-Home/。
原文摘要 · Abstract (English)
Computer-using agents powered by Vision-Language Models (VLMs) have demonstrated human-like capabilities in operating digital environments like mobile platforms. While these agents hold great promise for advancing digital automation, their potential for unsafe operations, such as system compromise and privacy leakage, is raising significant concerns. Detecting these safety concerns across the vast and complex operational space of mobile environments presents a formidable challenge that remains critically underexplored. To establish a foundation for mobile agent safety research, we introduce MobileRisk-Live, a dynamic sandbox environment accompanied by a safety detection benchmark comprising realistic trajectories with fine-grained annotations. Built upon this, we propose OS-Sentinel, a novel hybrid safety detection framework that synergistically combines a Formal Verifier for detecting explicit system-level violations with a VLM-based Contextual Judge for assessing contextual risks and agent actions. Experiments show that OS-Sentinel achieves 10%-30% improvements over existing approaches across multiple metrics. Further analysis provides critical insights that foster the development of safer and more reliable autonomous mobile agents. Our code and data are available at https://qiushisun.github.io/OS-Sentinel-Home/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。