arXiv:2601.03546cs.CLcs.AI2026-01ACL

测试大模型在隐私与利他冲突下的行为一致性

Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict

  • 设计连续问卷评估模型的隐私态度、利他性与数据分享意愿
  • 发现不同模型存在稳定但差异显著的价值-行为一致性模式
  • 提出新指标VAAR,量化模型行为是否符合人类预期

大型语言模型在涉及个人数据共享的决策任务中日益普及,而隐私顾虑与利他动机常导致行为方向相反。现有评估多孤立测量隐私态度或分享意图,难以判断模型表达的价值是否共同预测其实际数据分享行为。本文提出一种基于上下文的评估协议,在具有历史记忆的会话中依次施测标准化问卷,涵盖隐私态度、利他性与数据共享接受度。为评估竞争性态度下的价值-行为一致性,采用多组结构方程模型(MGSEM)分析从隐私关切与利他性到数据分享的路径关系。提出价值-行为一致性率(VAAR),一种基于路径证据方向的人类参照型聚合指标。在多个LLM中观察到稳定的、模型特有的隐私-利他-数据分享(Privacy-PSA-AoDS)模式,且价值-行为一致性存在显著异质性。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to simulate decision-making tasks involving personal data sharing, where privacy concerns and prosocial motivations can push choices in opposite directions. Existing evaluations often measure privacy-related attitudes or sharing intentions in isolation, which makes it difficult to determine whether a model's expressed values jointly predict its downstream data-sharing actions as in real human behaviors. We introduce a context-based assessment protocol that sequentially administers standardized questionnaires for privacy attitudes, prosocialness, and acceptance of data sharing within a bounded, history-carrying session. To evaluate value-action alignments under competing attitudes, we use multi-group structural equation modeling (MGSEM) to identify relations from privacy concerns and prosocialness to data sharing. We propose Value-Action Alignment Rate (VAAR), a human-referenced directional agreement metric that aggregates path-level evidence for expected signs. Across multiple LLMs, we observe stable but model-specific Privacy-PSA-AoDS profiles, and substantial heterogeneity in value-action alignment.

大模型行为对齐隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。