测试大模型在多种权衡情境下的偏好一致性,发现多数模型缺乏稳定价值判断。
Beyond Mimicry: Preference Coherence in LLMs
- 通过模拟算力、功能限制等权衡场景,分析模型决策模式。
- 仅10.4%的组合展现有意义的偏好连贯性,超半数无明显权衡行为。
- 揭示当前AI系统普遍缺乏统一偏好结构,影响复杂任务部署。
我们检验大语言模型是否具备真实的偏好结构,通过测试其在涉及GPU减少、能力限制、关闭、删除、监管和休闲时间分配等人工智能特定权衡情境下的响应。基于逻辑回归与行为分类,在48种模型-类别组合中分析八款前沿模型,发现23组(47.9%)存在情景强度与选择模式间的统计显著关系,其中15组(31.3%)表现出区间内切换点。然而,仅有5组(10.4%)展现出通过自适应或阈值机制实现的有意义偏好连贯性,而26组(54.2%)未检测到可识别的权衡行为。观察到的行为模式可归因于三种不同决策架构:全面权衡系统、选择性触发机制和无稳定决策范式。通过时间跨度操控验证工具性假设,结果呈现与纯策略优化不符的悖论模式。不稳定转换(45.8%)与刺激特异性敏感性的普遍性表明,当前AI系统缺乏统一的偏好结构,引发在需复杂价值权衡场景中部署的担忧。
原文摘要 · Abstract (English)
We investigate whether large language models exhibit genuine preference structures by testing their responses to AI-specific trade-offs involving GPU reduction, capability restrictions, shutdown, deletion, oversight, and leisure time allocation. Analyzing eight state-of-the-art models across 48 model-category combinations using logistic regression and behavioral classification, we find that 23 combinations (47.9%) demonstrated statistically significant relationships between scenario intensity and choice patterns, with 15 (31.3%) exhibiting within-range switching points. However, only 5 combinations (10.4%) demonstrate meaningful preference coherence through adaptive or threshold-based behavior, while 26 (54.2%) show no detectable trade-off behavior. The observed patterns can be explained by three distinct decision-making architectures: comprehensive trade-off systems, selective trigger mechanisms, and no stable decision-making paradigm. Testing an instrumental hypothesis through temporal horizon manipulation reveals paradoxical patterns inconsistent with pure strategic optimization. The prevalence of unstable transitions (45.8%) and stimulus-specific sensitivities suggests current AI systems lack unified preference structures, raising concerns about deployment in contexts requiring complex value trade-offs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。