arXiv:2606.31250cs.CLcs.AI2026-06

用AI评估文本风格抄袭,填补技术检测与欧盟版权法间的差距。

Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law

论文配图:Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law
图 1 · 摘自论文原文
  • 构建多维度评估框架PSALM,检测风格、叙事、内容等抽象相似性。
  • 微调后模型产生显著风格模仿,远超原文复制,涵盖情节模式等深层特征。
  • 适合法律合规、AI伦理研究者,推动技术与法律标准对接。

大规模语言模型(LLM)基于网络数据训练生成的内容可能构成版权侵权,但现有技术防护仅关注字面记忆。欧盟版权法采用更广义的“实质性相似”标准,涵盖风格选择、叙事结构和创作发挥。当前检测手段与法律保护范围不匹配,存在重大合规缺口。本文提出PSALM框架,通过十名评估者从计算重叠、风格(写作风格、叙述口吻)、内容(人物、情节、场景、世界观构建)及法定例外(戏仿、拼贴、引用、常见场景)等维度量化评估。在对Llama~3.2模型微调历史荷兰文学翻译数据集后发现:1)指令微调模型在接触语料前已存在非零基线风格相似性;2)微调引发系统性风格剽窃,覆盖所有侵权相关维度,超越字面复制,延伸至抽象叙事模式;3)负向偏好优化虽显著降低相似度,但仍残留可检测的风格痕迹。结果表明,仅针对文字复制的防护无法应对更广泛的版权风险。PSALM为可审计、法律导向的合规评估提供基础设施,但自动化相似度评分与侵权判定间的关系仍需法律专家验证。本工作弥合了定性法律标准与定量技术测量的鸿沟,揭示生成式AI与欧盟知识产权法之间的根本张力。

原文摘要 · Abstract (English)

Large language models (LLM) trained on web-scale corpora generate output that may infringe copyright, yet existing technical safeguards focus narrowly on verbatim memorisation. EU copyright doctrine applies a broader standards: substantial similarity, which extends to stylistic choices, narrative structure, and creative elaboration. This mismatch between what current methods detect and what the law protects leaves a significant compliance gap. We introduce PSALM, an LLM-as-a-judge framework that operationalises EU copyright doctrine through ten evaluators assessing computational overlap, stylistic dimensions (writing style, narrative voice), content dimensions (character, plot, scene, world building), and statutory exceptions (parody, pastiche, quotation, scènes à faire). Applying PSALM to Llama~3.2 models fine-tuned on translated historical Dutch literary works, we find that: 1) instruction-tuned models exhibit non-trivial baseline stylistic similarity prior to corpus exposure; 2) fine-tuning induces systematic stylistic appropriation across all infringement-relevant dimensions, extending beyond verbatim memorisation to abstract narrative patterns; 3) Negative Preference Optimisation unlearning substantially reduces similarity but leaves detectable residual stylistic patterns. These findings indicate that safeguards targeting literal copying alone are insufficient to mitigate broader copyright risks. PSALM provides infrastructure for auditable, legally informed compliance evaluation, though the relationship between automated similarity scores and infringement determinations requires validation by legal experts. This work bridges qualitative legal standards and quantitative technical measurement, exposing fundamental tensions between generative AI and EU intellectual property law.

版权评估风格抄袭AI合规法律科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。