为AI购物代理优化网站设计,提升其自主决策与执行效率。
Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability

- 构建三维度框架:可读性、可操作性、决策可靠性
- 代理成功率从49.3%升至89.3%,步骤减少约30%
- 适合开发电商系统或研究AI代理交互的团队
在线购物正转向由AI代理自主搜索商品、比较选项、评估约束并完成部分购买流程。网站设计需同时支持人类与代理交互。本文提出“代理就绪网站”框架,提升电商平台对AI代理的可读性、可解释性、可验证性和可操作性。现有网页设计、SEO及生成引擎优化(GEO)指标无法充分评估网站对代理交互的支持能力。该框架围绕三个维度——代理可解释性、代理可执行性、代理决策可靠性——通过机器可读性、语义清晰度、代理可操作性及上下文决策可靠性信号实现。在控制实验中,对比相同原型网站的人类导向基线与代理就绪版本,包含五个任务、三种浏览器代理模型(GPT-4.1、Gemini-2.5 Flash、Grok-4 Fast)和300次运行。结果显示,代理就绪网站达成134次成功(总150次),基线仅74次(严格成功率89.3% vs. 49.3%),尤其在商品详情提取、比价和多约束选择上提升显著;部分成功从43降至3,平均步骤数由9.31降至6.49。结果表明,结构清晰度、操作提示、证据信号与时间有效性标识能显著提升AI浏览器代理的可靠性和效率。
原文摘要 · Abstract (English)
Online shopping is increasingly shifting toward a model in which AI agents independently search for products, compare options, evaluate constraints, and carry out parts of the purchasing process for users. Website design must now support both human and agent-mediated interaction. This paper introduces the agent-ready website, a design framework for enhancing the readability, interpretability, verifiability, and actionability of e-commerce platforms for AI agents. Existing web design, SEO, and generative engine optimization (GEO) metrics do not fully assess a website's capacity for agent-mediated interaction. The proposed framework is structured around three dimensions agent interpretability, agent executability, and agent decision reliability supported by features such as machine readability, semantic clarity, agent actionability, and contextual decision-reliability signals. The framework is evaluated through a controlled experiment comparing a human-oriented baseline and an agent-ready version of an identical website prototype, with identical catalogs, pricing, stock, and shopping workflows. The evaluation involved five tasks, three browser-agent models (GPT-4.1, Gemini-2.5 Flash, and Grok-4 Fast), and 300 runs, measuring PASS,PARTIAL,FAIL outcomes, strict and functional success rates, error patterns, step counts, and token consumption. The agent-ready website achieved 134 PASS runs out of 150 versus 74 out of 150 for the baseline (strict success rates of 89.3% vs. 49.3%), with the largest gains in product detail extraction, comparison, and multi-constraint selection. It also reduced PARTIAL outcomes from 43 to 3 and lowered the average step count from 9.31 to 6.49. These results provide preliminary evidence that enhanced structural clarity, action cues, evidence signals, and temporal validity indicators can substantially improve the reliability and efficiency of AI browser agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。