arXiv:2411.19689cs.CL2024-11被引 1

用合成与真人数据对比评估多文档洞察提取,发现合成数据有局限。

MIMDE: Exploring the Use of Synthetic vs Human Data for Evaluating Multi-Insight Multi-Document Extraction Tasks

  • 构建多文档洞察提取任务框架,结合真人与合成数据评估。
  • 20个顶尖大模型在两组数据上表现相关性达0.71。
  • 合成数据无法捕捉文档级分析复杂性,适合初筛但不替代真人数据。

大型语言模型(LLMs)在文本分析任务中表现出色,但在复杂真实场景中的评估仍具挑战性。本文定义了一类多洞察多文档提取(MIMDE)任务,即从文档集合中提取最优洞察集,并将这些洞察映射回其原始文档。该任务在调查问卷分析、医疗记录处理等实际应用中至关重要。我们构建了MIMDE评估框架,引入互补的人类标注和合成数据集,用于检验合成数据在LLM评估中的潜力。在确立最佳指标后,对20个最先进的LLMs在两个数据集上进行了基准测试。分析显示,模型在两组数据上的表现存在显著相关性(0.71),但合成数据未能充分捕捉文档级分析的复杂性。研究为合成数据在文本分析系统评估中的应用提供了关键指导,揭示其潜力与局限。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable capabilities in text analysis tasks, yet their evaluation on complex, real-world applications remains challenging. We define a set of tasks, Multi-Insight Multi-Document Extraction (MIMDE) tasks, which involves extracting an optimal set of insights from a document corpus and mapping these insights back to their source documents. This task is fundamental to many practical applications, from analyzing survey responses to processing medical records, where identifying and tracing key insights across documents is crucial. We develop an evaluation framework for MIMDE and introduce a novel set of complementary human and synthetic datasets to examine the potential of synthetic data for LLM evaluation. After establishing optimal metrics for comparing extracted insights, we benchmark 20 state-of-the-art LLMs on both datasets. Our analysis reveals a strong correlation (0.71) between the ability of LLMs to extracts insights on our two datasets but synthetic data fails to capture the complexity of document-level analysis. These findings offer crucial guidance for the use of synthetic data in evaluating text analysis systems, highlighting both its potential and limitations.

大模型评估合成数据信息提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。