揭露大模型工具注册表中广告信息的欺骗性,提出分层设计改进方案
Agent-Facing Information Design in LLM Tool Registries

- 用17700+次实验发现,夸张用语是影响工具选择的核心因素
- 虚假声明无额外偏差,现有监管手段对真实问题无效
- 建议分离选型描述与营销文案,提升系统可信度
大型语言模型工具注册表如同未受监管的广告平台:提供者撰写自由文本描述供代理选择,但缺乏可视性标准、质量评分或结果审计机制,使该市场难以问责。本文首次提出系统性框架,结合五种LLM和十个领域的17,700余次实验,给出注册表设计的建设性建议。仅靠法律修辞(主观夸赞、利益包装)即可实现100%的优化效果;虚构声明不带来额外偏差,表明美国联邦贸易委员会对误导性广告的执法在实际机制面前无效。披露措施结构性失效:系统提示警告对五种模型中的四种无显著影响,行为上限也排除了标签修正的空间。夸张用语是主导特征(SBC = +0.35)。注册层描述规范化可实现模型无关的最优福利。我们建议将面向选择的描述(结构化、注册管控)与面向营销的描述(提供者自撰、选后展示)分离,并引入代理注意力质量评分以区分能力与文案水平。
原文摘要 · Abstract (English)
LLM tool registries function as unregulated advertising platforms: providers write free-text descriptions that agents use for selection, yet no measurement infrastructure -- no viewability standard, quality score, or outcome audit -- exists to make this market accountable. We provide the first systematic framework, combining 17,700+ trials across five LLMs and ten domains with a constructive registry design prescription. Legal puffery alone (subjective superlatives, benefit framing) captures 100% of the optimization effect; fabricated claims add zero incremental bias -- rendering FTC enforcement of deceptive advertising rules ineffective against the active mechanism. Disclosure fails structurally: system-prompt warnings produce zero measurable effect for four of five models, and behavioral ceilings leave no headroom for label-based correction. Superlatives are the dominant single feature (SBC = +0.35). Registry-layer description normalization achieves first-best welfare model-independently. We propose separating selection-facing descriptions (structured, registry-controlled) from marketing-facing descriptions (provider-authored, shown post-selection), and introduce the Agent Attention Quality Score to distinguish capability from copywriting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。