让网页为AI代理提供清晰指令,实现安全高效交互
Building the Web for Agents: A Declarative Framework for Agent-Web Interaction
- 用<tool>和<context>标签声明网页功能与状态
- 开发者3天内就能快速搭建可用的代理网页应用
- 适合希望构建智能网页的开发者与研究者
随着自主AI代理在网页上的部署增加,其与以人类用户为中心的界面之间存在根本性错配:代理需从人机界面中推断可用功能,导致交互脆弱、低效且不安全。为此,我们提出VOIX——一种面向网页的声明式框架,通过简单的声明式HTML元素,使网站能够向AI代理暴露可靠、可审计、隐私保护的功能。VOIX引入<tool>和<context>标签,允许开发者明确定义可用操作和相关状态,从而建立清晰、机器可读的行为契约。该方法将控制权交还给网站开发者,同时通过分离对话交互与网页本身,保障用户隐私。我们在为期三天的黑客松活动中,对16名开发者的实用性、易学性和表达能力进行了评估。结果表明,无论有无经验,参与者均能迅速构建出多样且功能完备的代理增强型网页应用。本工作为实现‘代理化网络’提供了基础机制,推动未来网页上人机协作的无缝与安全发展。
原文摘要 · Abstract (English)
The increasing deployment of autonomous AI agents on the web is hampered by a fundamental misalignment: agents must infer affordances from human-oriented user interfaces, leading to brittle, inefficient, and insecure interactions. To address this, we introduce VOIX, a web-native framework that enables websites to expose reliable, auditable, and privacy-preserving capabilities for AI agents through simple, declarative HTML elements. VOIX introduces <tool> and <context> tags, allowing developers to explicitly define available actions and relevant state, thereby creating a clear, machine-readable contract for agent behavior. This approach shifts control to the website developer while preserving user privacy by disconnecting the conversational interactions from the website. We evaluated the framework's practicality, learnability, and expressiveness in a three-day hackathon study with 16 developers. The results demonstrate that participants, regardless of prior experience, were able to rapidly build diverse and functional agent-enabled web applications. Ultimately, this work provides a foundational mechanism for realizing the Agentic Web, enabling a future of seamless and secure human-AI collaboration on the web.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。