arXiv:2409.07355cs.CL2024-09被引 26

人机协作生成文本评估属性,提升检查表式评价效果。

Think Together and Work Better: Combining Humans' and LLMs' Think-Aloud Outcomes for Effective Text Evaluation

  • 融合人类思考与大模型推理,通过听觉思维法生成评估维度。
  • 在连贯性、流畅性、一致性、相关性四维上均优于传统方法。
  • 适合需要高质量文本评估的AI研发与内容审核场景。

本研究提出InteractEval框架,结合人类专家与大语言模型(LLMs),采用听觉思维(Think-Aloud, TA)方法生成检查表式文本评估的属性。通过融合人类灵活性与推理能力及大模型的一致性,InteractEval在连贯性、流畅性、一致性和相关性四个维度上均超越传统非大模型与纯大模型基线。实验表明,TA方法能促进人类与大模型的发散思维,生成更广泛的相关属性,从而提升评估性能。对比分析显示,人类在识别内部质量属性(连贯性与流畅性)上表现更优,而大模型在外部对齐属性(一致性与相关性)上更具优势。因此,联合利用人与大模型可获得最佳评估结果。本研究强调在自动化检查表式文本评估中有效整合人机协作的必要性。代码已公开于https://github.com/BBeeChu/InteractEval.git。

原文摘要 · Abstract (English)

This study introduces \textbf{InteractEval}, a framework that integrates human expertise and Large Language Models (LLMs) using the Think-Aloud (TA) method to generate attributes for checklist-based text evaluation. By combining human flexibility and reasoning with LLM consistency, InteractEval outperforms traditional non-LLM-based and LLM-based baselines across four distinct dimensions, consisting of Coherence, Fluency, Consistency, and Relevance. The experiment also investigates the effectiveness of the TA method, showing that it promotes divergent thinking in both humans and LLMs, leading to the generation of a wider range of relevant attributes and enhance text evaluation performance. Comparative analysis reveals that humans excel at identifying attributes related to internal quality (Coherence and Fluency), but LLMs perform better at those attributes related to external alignment (Consistency and Relevance). Consequently, leveraging both humans and LLMs together produces the best evaluation outcomes. In other words, this study emphasizes the necessity of effectively combining humans and LLMs in an automated checklist-based text evaluation framework. The code is available at \textbf{\url{https://github.com/BBeeChu/InteractEval.git}}.

文本评估人机协作大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。