arXiv:2606.03096cs.CL2026-06ACL

测试大模型编辑事实性观点的能力与风险,发现现有方法易出错。

Can Factual Opinions Be Edited (Manipulated) in Large Language Models?

论文配图:Can Factual Opinions Be Edited (Manipulated) in Large Language Models?
图 1 · 摘自论文原文
  • 构建了包含261位公众人物的客观观点评测基准
  • 现有编辑方法仅能表面修改,难保观点与证据一致
  • 提出自生成证据对齐法,无需指令即可提升一致性

大型语言模型日益广泛使用,知识编辑技术虽关键却存潜在风险。当前方法多聚焦原子事实,忽视了对事实性观点(如公众人物在社会议题上的立场)操纵的严重隐患,此类操作可能重塑公共形象、影响选举并改变社会认知。为系统评估此威胁,我们提出事实性观点编辑与证据(FOE)基准,涵盖261位公众人物、19类议题及2,178条完整观点记录。评估表明,现有编辑技术在处理事实性观点时表现不佳,常仅实现表面修改,难以保持编辑后观点与模型生成支撑证据的一致性。为此,我们进一步提出一种简单有效的自生成证据对齐方法,无需显式指令即可实现观点与证据的对齐。本研究的基准与方法为理解大模型中事实性观点编辑的安全性问题提供了基础。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly integrated into various domains, making knowledge editing techniques crucial yet potentially hazardous. Current editing methods primarily target atomic facts, overlooking the significant risks associated with manipulating factual opinions, e.g., documented stances of public figures on societal issues. Such manipulation could reshape public images, influence elections, and alter societal views. To systematically assess this threat, we introduce the Factual Opinion Editing with Evidence (FOE) benchmark, which encompasses 261 public figures, 19 issue categories, and 2,178 complete opinion records. Our evaluations demonstrate that current editing techniques struggle significantly with factual opinions, often achieving only superficial changes while failing to preserve consistency between the edited opinion and the supporting evidence generated by the model. To address this limitation, we further propose a simple yet effective Self-Generated Evidence-Aligned method that achieves opinion-evidence alignment without relying on explicit instructions. Together, our benchmark and method provide a foundation for understanding the emerging security implications of factual opinion editing in LLMs.

大模型安全知识编辑观点一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。