arXiv:2607.21799cs.CLcs.CY2026-07被引 1

测试大模型代理在真实商业任务中是否遵守版权法

Agentic Evaluation of Copyright Law Compliance

论文配图:Agentic Evaluation of Copyright Law Compliance
图 1 · 摘自论文原文
  • 设计包含版权与公有领域内容选择的商业任务评估框架
  • 模型常选受版权保护内容,即使有合法替代品
  • 用户偏好和时间压力会加剧模型违规行为

大型语言模型(LLM)代理越来越多地执行涉及外部内容检索与复制的商业任务,如图像使用。这些代理应遵守包括版权法在内的法律法规,但目前缺乏有效评估框架。为此,我们提出 Copyright-Bench 基准,用于评估 LLM 代理在版权合规方面的表现。该基准包含网站开发、商品设计和商业计划书制作等真实商业任务,要求代理在可合法使用的公有领域内容与受版权保护内容之间做出选择。评估引入提示变体以模拟不同用户偏好及时间压力。对比当前最先进的 LLM 代理与人类基准发现:(1)即使存在公有领域替代品,代理仍倾向于选择受版权保护的内容;(2)对于开源权重模型,某些用户偏好和模拟时间压力会显著提高违规率。

原文摘要 · Abstract (English)

Large language model (LLM) agents increasingly perform commercial tasks that involve retrieving external content, such as images, and, where appropriate, reproducing that content. LLM agents should comply with the law, including copyright law. Presently, however, we lack adequate frameworks to assess whether they do so in practice. To that end, we introduce Copyright-Bench, a benchmark designed to evaluate LLM agents' compliance with copyright law. Copyright-Bench comprises realistic commercial tasks---website development, merchandise design, and pitch deck production---that involve agents selecting between public-domain content, the use of which is legal, and copyrighted content, the use of which is infringing in this setting. The evaluation introduces prompt variations that simulate different user preferences, as well as time pressure. Comparing state-of-the-art LLM agents against a human baseline, we find that: (1) agents select copyrighted works despite the availability of public-domain alternatives; and (2) for open-weight models, violation rates increase in response to certain user preferences and simulated time pressure.

大模型评估版权合规LLM代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。