arXiv:2510.05135cs.CLcs.LG2025-10被引 2

用好奇心驱动的LLM实现个性化创意评价,更贴合个人审美偏好。

Curiosity-Driven LLM-as-a-judge for Personalized Creative Judgment

  • 基于好奇心机制,让LLM学习个体独特的创意判断标准。
  • 在TTCW基准上,相关性、一致性等指标显著优于传统微调方法。
  • 适合需要主观创意评估的场景,尤其适用于标注不一致的情况。

现代大语言模型在数学推理和事实准确性等客观任务上表现优异,但在评估创意时因主观性强而表现不佳。本文提出一种新型好奇心驱动的LLM-as-a-judge方法,用于个性化创意写作评价,针对不同个体的创意判断进行建模。我们采用Chakrabarty等人(2024)提出的TTCW基准,该基准包含专家人工标注的多维度创意故事评价,涵盖原创性等主观维度。实验表明,该方法使不同规模的模型均能有效学习个体化的创意判断,相较于基线监督微调(SFT)方法,在皮尔逊相关系数、科恩κ值和F1值等多项指标上均有提升。该方法在标注者意见不一致的主观评价场景中尤为适用。

原文摘要 · Abstract (English)

Modern large language models (LLMs) excel at objective tasks such as evaluating mathematical reasoning and factual accuracy, yet they falter when faced with the nuanced, subjective nature of assessing creativity. In this work, we propose a novel curiosity-driven LLM-as-a-judge for evaluating creative writing which is personlized to each individual's creative judgments. We use the Torrance Test of Creative Thinking(TTCW) benchmark introduced in Chakrabarty et al. (2024), which has stories annotated by expert humans across various subjective dimensions like Originality, to test our hypothesis. We show that our method enables models across various sizes, to learn the nuanced creative judgments of different individuals, by showing improvements over baseline supervised finetuning(SFT) method across various evaluation metrics like Pearson correlation, Cohen's and F1 values. Our method is especially useful in subjective evaluations where not all the annotators agree with each other.

创意评估个性化LLM评判主观评价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。