arXiv:2509.21267cs.CLcs.CY2025-09被引 7

区分任务类型,让大模型在该变时变、该稳时稳,避免无意义重复。

Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework

  • 按任务设计多样性标准,判断输出是否真正不同。
  • 用户研究验证了分类体系符合人类对差异的感知。
  • 仅在需要多样性时增强差异,不牺牲质量。

大型语言模型常产生同质化输出,但是否构成问题取决于具体任务。对于客观数学任务,解题策略可异但答案必须一致;而创意写作中,情节、设定等核心要素应有变化,不仅限于词汇多样性。现有研究很少基于任务定义多样性。本文提出四项贡献:(1)构建任务分类体系,明确不同任务下的功能多样性标准——用户是否会认为两个响应在任务上存在实质差异;(2)通过小型用户研究验证该分类与人类对功能多样性的感知一致;(3)设计任务依赖的采样方法,仅在不希望同质化时提升多样性;(4)提供证据反驳普遍存在的‘多样性-质量权衡’假说,指出其源于对多样性与质量在任务无关视角下的误判。

原文摘要 · Abstract (English)

Large language models often generate homogeneous outputs, but whether this is problematic depends on the specific task. For objective math tasks, responses may vary in terms of problem-solving strategy but should maintain the same verifiable answer. Whereas, for creative writing tasks, we often expect variation in key narrative components (e.g. plot, setting, etc.) beyond mere vocabulary diversity. Prior work on homogenization rarely conceptualizes diversity in a task-dependent way. We address this gap with four contributions: (1) a task taxonomy with distinct notions of functional diversity -- whether a user would perceive two responses as meaningfully different for a given task; (2) a small user study validating that the taxonomy aligns with human perception of functional diversity; (3) a task-dependent sampling technique that increases diversity only where homogenization is undesired; (4) evidence challenging the perceived diversity-quality trade-off, showing it may stem from mis-conceptualizing both diversity and quality in a task-agnostic way.

大模型多样性评估任务分类输出优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。