用四类目标框架解析大模型输出多样性,揭示优化一个目标可能损害其他方面。
Magic, Madness, Heaven, Sin: LLM Output Diversity is Everything, Everywhere, All at Once
- 提出魔、疯、天、罪四类范式,按任务目标分类输出变异
- 发现安全优化会意外降低代表性与创意多样性
- 适合关注评估公平性与多目标权衡的研究者
大型语言模型研究常关注生成、推理、对齐和表征分析中的输出差异,统称为‘多样性’。然而术语分散,因任务的规范目标常未明确。本文提出‘魔、疯、天、罪’框架,将输出变异置于同质-异质轴上,由任务及其规范目标决定价值。将任务分为四类:认知(事实性)、交互(用户效用)、社会(代表性)与安全(鲁棒性)。针对每类分析其失效模式与术语如幻觉、模式坍缩、偏见与抹除。通过分析所有跨情境交互,发现优化某一目标(如安全性)可能无意损害群体代表性或创造性多样性。主张应基于上下文评估输出变异,将其视为由任务目标塑造的属性,而非模型固有特性。
原文摘要 · Abstract (English)
Research on Large Language Models (LLMs) studies output variation across generation, reasoning, alignment, and representational analysis, often under the umbrella of "diversity." Yet the terminology remains fragmented, largely because the normative objectives underlying tasks are rarely made explicit. We introduce the Magic, Madness, Heaven, Sin framework, which models output variation along a homogeneity-heterogeneity axis, where valuation is determined by the task and its normative objective. We organize tasks into four normative contexts: epistemic (factuality), interactional (user utility), societal (representation), and safety (robustness). For each, we examine the failure modes and vocabulary such as hallucination, mode collapse, bias, and erasure through which variation is studied. We apply the framework to analyze all pairwise cross-contextual interactions, revealing that optimizing for one objective, such as improving safety, can inadvertently harm demographic representation or creative diversity. We argue for context-aware evaluation of output variation, reframing it as a property shaped by task objectives rather than a model's intrinsic trait.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。