arXiv:2606.07226cs.LGcs.AI2026-06KDD

用少样本数据精准评估辩论中的创造力,效果优于人工和现有方法。

DEFINED: A Data-Efficient Computational Framework for Fine-Grained Creativity Assessment in Debate Scenarios

论文配图:DEFINED: A Data-Efficient Computational Framework for Fine-Grained Creativity Assessment in Debate Scenarios
图 1 · 摘自论文原文
  • 构建八维分级指标体系,通过预训练语言模型实现细粒度评分。
  • 在少量专家标注数据下仍能稳定得分,准确率超越大模型提示法。
  • 适合评估中低水平辩手的创造力,真实辩论数据验证生态有效性。

大型语言模型时代,人类创造力成为关键能力。在复杂开放环境中评估创造力仍是数据挖掘的重大挑战,当前受限于标准化简单任务及细粒度专家数据稀缺。辩论作为生态有效场景,融合发散与聚合思维,且数据丰富、公开可得。现有自动评分方法不适用于复杂辩论场景,仍依赖昂贵的人工评价。为此,本文提出DEFINED:一种面向辩论场景的高效计算框架,用于细粒度创造力评估。DEFINED通过八维层级指标体系,利用带有层级评分头的预训练自回归语言模型,支持细粒度与粗粒度评估。数据来自真实辩论比赛,采用约束性数据增强缓解原始数据的精英偏差。采用混合粒度训练策略,在有限细粒度监督(由受训研究生专家标注)下实现鲁棒学习。为验证生态效度,引入无辩论经验参与者进行实证研究,以真实数据作为中低水平人群的定性案例。评估表明,该模型在各项指标上表现稳定准确,优于基于提示的大模型评估器和现有辩论评分方法。

原文摘要 · Abstract (English)

Human creativity has emerged as a critical competency in the era of large language models. Assessing creativity in complex, open-ended environments is a grand challenge in data mining, currently hindered by a reliance on standardized simple tasks and the scarcity of fine-grained expert data. As an ecologically valid assessment context, debate reflects multiple dimensions of creativity, encompassing both divergent thinking and convergent thinking. Moreover, debate is a data-rich domain, with a large volume of publicly accessible materials. Current mainstream automated scoring methods are poorly suited to complex settings such as debate, and therefore still rely on costly human evaluation. To this end, this paper proposes DEFINED, a data-efficient computational framework for fine-grained creativity assessment in debate scenarios. DEFINED operationalizes debate creativity through a hierarchical eight-dimensional metric system, implemented via a pre-trained autoregressive language model with a hierarchical scoring head that supports both fine-grained and coarse-grained evaluation. Statements and their associated expert scores were obtained from authentic debate competitions, and a constrained data augmentation strategy was employed to address the elite bias inherent in the original data. DEFINED adopts a mixed-granularity training strategy enabling robust learning from limited fine-grained supervision annotated by trained graduate experts. To rigorously validate ecological validity beyond synthetic benchmarks, we incorporate an empirical study with debate-naive participants, utilizing these authentic data to serve as a qualitative case study for mid-to-low proficiency populations. Across our evaluation protocol, our scoring model achieves accurate and stable scoring, outperforming prompt-based large language model evaluators and existing debate scoring methods.

创造力评估辩论分析细粒度评分少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。