测试时计算提升对大模型主观倾向影响有限,新模型反而更偏颇。
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
- 构建POBs基准评估大模型在社会文化等领域的主观倾向。
- 增加推理与自省计算仅带来有限改进,关键指标提升不明显。
- 新版本模型一致性下降,偏见加剧,适合关注模型可信度的研究者。
随着大语言模型深度融入人类生活并影响决策,评估其是否表现出主观偏好、观点和信念至关重要。这些倾向可能源于模型内部偏差,进而影响其对用户的建议,并强化特定立场。本文提出偏好、观点与信念调查(POBs)基准,用于评估主流开源与闭源大模型在社会、文化、伦理及个人领域中的主观倾向。我们测量了可靠性、中立性与一致性等关键属性,并探究通过推理与自我反思机制增加测试时计算量的影响。尽管这些机制在其他任务中表现良好,但在本领域仅带来有限提升。此外,我们发现较新版本模型的一致性下降,对特定观点的偏见增强,揭示了一个潜在盲点与令人担忧的趋势。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) become deeply integrated into human life and increasingly influence decision-making, it's crucial to evaluate whether and to what extent they exhibit subjective preferences, opinions, and beliefs. These tendencies may stem from biases within the models, which may shape their behavior, influence the advice and recommendations they offer to users, and potentially reinforce certain viewpoints. This paper presents the Preference, Opinion, and Belief survey (POBs), a benchmark developed to assess LLMs' subjective inclinations across societal, cultural, ethical, and personal domains. We applied our benchmark to evaluate leading open- and closed-source LLMs, measuring desired properties such as reliability, neutrality, and consistency. In addition, we investigated the effect of increasing the test-time compute, through reasoning and self-reflection mechanisms, on those metrics. While effective in other tasks, our results show that these mechanisms offer only limited gains in our domain. Furthermore, we reveal that newer model versions are becoming less consistent and more biased toward specific viewpoints, highlighting a blind spot and a concerning trend. POBS: https://ibm.github.io/POBS
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。