arXiv:2504.12344cs.CL2025-04中稿 · ICLR被引 3

发现大模型对敏感概念易产生立场化偏移,提出可检测的审计方法。

Propaganda AI: An Analysis of Semantic Divergence in Large Language Models

  • 通过语义熵与跨模型分歧识别模型异常响应
  • 在12个敏感话题中9个发现模型特异性偏移
  • 适合模型安全评估与部署后监控使用

大型语言模型可能表现出概念条件下的语义偏移:常见高层线索(如意识形态、公众人物)会引发异常一致的立场化回应,且逃避令牌触发审计。这种行为目前安全评估存在盲区,但影响广泛,因这些概念线索可大规模引导内容呈现。本文提出RAVEN(Response Anomaly Vigilance)黑盒审计方法,结合改写样本的语义熵与跨模型分歧,标记同时高置信度且异于同侪的情况。在控制的LoRA微调实验中,仅用小规模偏见语料即可植入概念立场。对五类大模型在十二个敏感话题(每模型360个提示)进行审计,基于双向蕴含聚类,发现9个话题存在模型特异性偏移。概念级审计补充了令牌级防护,为发布评估与部署后监控提供实用预警信号。

原文摘要 · Abstract (English)

Large language models (LLMs) can exhibit concept-conditioned semantic divergence: common high-level cues (e.g., ideologies, public figures) elicit unusually uniform, stance-like responses that evade token-trigger audits. This behavior falls in a blind spot of current safety evaluations, yet carries major societal stakes, as such concept cues can steer content exposure at scale. We formalize this phenomenon and present RAVEN (Response Anomaly Vigilance), a black-box audit that flags cases where a model is simultaneously highly certain and atypical among peers by coupling semantic entropy over paraphrastic samples with cross-model disagreement. In a controlled LoRA fine-tuning study, we implant a concept-conditioned stance using a small biased corpus, demonstrating feasibility without rare token triggers. Auditing five LLM families across twelve sensitive topics (360 prompts per model) and clustering via bidirectional entailment, RAVEN surfaces recurrent, model-specific divergences in 9/12 topics. Concept-level audits complement token-level defenses and provide a practical early-warning signal for release evaluation and post-deployment monitoring against propaganda-like influence.

大模型安全语义偏移审计工具宣传风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。