LLM在社区对不确定性态度上表现出深层行为一致性,超越简单记忆。
Community-Aligned Behavior Under Uncertainty: Evidence of Epistemic Stance Transfer in LLMs
- 通过删除关键事件知识,测试模型在未知下的行为模式
- 移除事实后仍保持社区特异的应对不确定性方式
- 适合关注模型偏见与安全部署的研究者
当大语言模型(LLMs)被对齐至特定在线社区时,它们是否展现出可泛化的、反映该社区态度和应对新不确定性的行为模式,还是仅在复现训练数据中的模式?我们提出一种框架,通过有针对性地删除事件知识,并用多重探测验证,评估模型在信息缺失下是否仍能再现社区的自然响应模式。基于俄乌军事话语和美国党派推特数据,发现即使经过激进的事实清除,对齐后的模型仍维持稳定且具有社区特异性的不确定性处理行为。结果表明,对齐不仅包含表面模仿,还编码了结构化、可泛化的行为特征。该框架为系统检测模型在无知状态下持续存在的行为偏差提供了方法,有助于推动更安全、透明的大模型应用。
原文摘要 · Abstract (English)
When large language models (LLMs) are aligned to a specific online community, do they exhibit generalizable behavioral patterns that mirror that community's attitudes and responses to new uncertainty, or are they simply recalling patterns from training data? We introduce a framework to test epistemic stance transfer: targeted deletion of event knowledge, validated with multiple probes, followed by evaluation of whether models still reproduce the community's organic response patterns under ignorance. Using Russian--Ukrainian military discourse and U.S. partisan Twitter data, we find that even after aggressive fact removal, aligned LLMs maintain stable, community-specific behavioral patterns for handling uncertainty. These results provide evidence that alignment encodes structured, generalizable behaviors beyond surface mimicry. Our framework offers a systematic way to detect behavioral biases that persist under ignorance, advancing efforts toward safer and more transparent LLM deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。