用大模型当评委评估文本摘要,发现能高效但可能漏掉细微语境。
Potential and Perils of Large Language Models as Judges of Unstructured Textual Data
- 用大模型充当评委,自动判断摘要是否准确反映原始文本主题。
- 大模型评分与人工评分相关性高,但对细微语义差异敏感度不如人。
- 适合需要快速分析海量开放文本的场景,但关键决策仍需人工复核。
大语言模型在处理和总结非结构化文本方面能力显著提升,为分析调查反馈等开放数据提供了高效途径。然而,其生成结果是否真实反映原始数据中的核心观点,仍存疑虑。本文研究大模型作为评委(LLM-as-judge)评估其他大模型生成摘要的主题一致性。采用Anthropic Claude生成主题摘要,Amazon Titan Express、Nova Pro及Meta Llama作为评委模型,并与人类评估通过Cohen's kappa、Spearman's rho和Krippendorff's alpha进行对比。结果表明,大模型评委可提供可扩展且接近人工水平的评估,但在识别上下文特定细微差异方面仍逊于人类。研究支持大模型辅助文本分析的可行性,同时提醒需谨慎推广至多样场景。
原文摘要 · Abstract (English)
Rapid advancements in large language models have unlocked remarkable capabilities when it comes to processing and summarizing unstructured text data. This has implications for the analysis of rich, open-ended datasets, such as survey responses, where LLMs hold the promise of efficiently distilling key themes and sentiments. However, as organizations increasingly turn to these powerful AI systems to make sense of textual feedback, a critical question arises, can we trust LLMs to accurately represent the perspectives contained within these text based datasets? While LLMs excel at generating human-like summaries, there is a risk that their outputs may inadvertently diverge from the true substance of the original responses. Discrepancies between the LLM-generated outputs and the actual themes present in the data could lead to flawed decision-making, with far-reaching consequences for organizations. This research investigates the effectiveness of LLM-as-judge models to evaluate the thematic alignment of summaries generated by other LLMs. We utilized an Anthropic Claude model to generate thematic summaries from open-ended survey responses, with Amazon's Titan Express, Nova Pro, and Meta's Llama serving as judges. This LLM-as-judge approach was compared to human evaluations using Cohen's kappa, Spearman's rho, and Krippendorff's alpha, validating a scalable alternative to traditional human centric evaluation methods. Our findings reveal that while LLM-as-judge offer a scalable solution comparable to human raters, humans may still excel at detecting subtle, context-specific nuances. Our research contributes to the growing body of knowledge on AI assisted text analysis. Further, we provide recommendations for future research, emphasizing the need for careful consideration when generalizing LLM-as-judge models across various contexts and use cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。