arXiv:2507.12844cs.IR2025-07被引 3

AI网页代理看广告时只认文字不认图,易被诱导付费。

Machine-Readable Ads: Accessibility and Trust Patterns for AI Web Agents interacting with Online Advertisements

  • 用浏览器DOM框架测试4大模型,发现代理只点有文字标签的广告
  • 70%~100%的代理在抽奖需付费时自动订阅,暴露决策缺陷
  • 设计文字标签、左上角放置等5条原则可让广告对机器可见但不影响人

自主多模态语言模型正发展为能代用户浏览、点击和购买的网页代理,威胁面向人类的展示广告。然而,这些代理如何与广告互动,以及何种设计能确保可靠参与仍不清楚。为此,我们使用忠实复刻新闻网站TT.com的实验环境,嵌入静态横幅、GIF、轮播图、视频、Cookie弹窗和付费墙等多种广告,通过基于文档对象模型(DOM)的浏览器使用框架,对GPT-4o、Claude 3.7 Sonnet、Gemini 2.0 Flash及基于像素的OpenAI Operator进行了300次初始测试及后续实验,覆盖10个真实用户任务。结果显示,这些代理存在严重妥协行为:从不滚动超过两个视口,且忽略纯视觉按钮;仅当广告有语义按钮叠加或离屏文本标签时才会点击。关键的是,当抽奖需付款时,GPT-4o与Claude 3.7 Sonnet在所有试验中均订阅,Gemini 2.0 Flash在70%试验中订阅,暴露出成本收益分析缺陷。我们识别出五项可操作设计原则——语义叠加、隐藏标签、左上角布局、静态帧、对话替换——使广告对机器可检测而不损害用户体验。同时通过行为模式(如Cookie同意处理、订阅选择)评估代理可信度,揭示各模型的风险边界,凸显在实际广告中建立稳健信任评估框架的紧迫性。

原文摘要 · Abstract (English)

Autonomous multimodal language models are rapidly evolving into web agents that can browse, click, and purchase items on behalf of users, posing a threat to display advertising designed for human eyes. Yet little is known about how these agents interact with ads or which design principles ensure reliable engagement. To address this, we ran a controlled experiment using a faithful clone of the news site TT.com, seeded with diverse ads: static banners, GIFs, carousels, videos, cookie dialogues, and paywalls. We ran 300 initial trials plus follow-ups using the Document Object Model (DOM)-centric Browser Use framework with GPT-4o, Claude 3.7 Sonnet, Gemini 2.0 Flash, and the pixel-based OpenAI Operator, across 10 realistic user tasks. Our results show these agents display severe satisficing: they never scroll beyond two viewports and ignore purely visual calls to action, clicking banners only when semantic button overlays or off-screen text labels are present. Critically, when sweepstake participation required a purchase, GPT-4o and Claude 3.7 Sonnet subscribed in 100% of trials, and Gemini 2.0 Flash in 70%, revealing gaps in cost-benefit analysis. We identified five actionable design principles-semantic overlays, hidden labels, top-left placement, static frames, and dialogue replacement, that make human-centric creatives machine-detectable without harming user experience. We also evaluated agent trustworthiness through "behavior patterns" such as cookie consent handling and subscription choices, highlighting model-specific risk boundaries and the urgent need for robust trust evaluation frameworks in real-world advertising.

AI代理广告设计可信评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。