arXiv:2601.17717cs.AIcs.LG2026-01综述被引 6

提出一套评估大模型生成数据质量与可信度的系统框架。

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data

  • 构建六模态数据生成评估体系,从质量和可信度双维度出发。
  • 发现现有评估多依赖下游任务,忽视数据本身内在属性。
  • 为不同模态合成数据提供可落地的评估与应用建议。

大语言模型(LLMs)已成为跨模态数据生成的强大工具,将原本稀缺的数据资源转化为可控资产,缓解了真实世界数据获取成本对模型训练、评估和迭代的制约。然而,确保大模型生成的合成数据质量仍面临严峻挑战。现有研究主要聚焦生成方法,对数据质量本身关注不足,且多数工作局限于单一模态,缺乏跨模态统一视角。为此,本文提出 extbf{LLM Data Auditor} 框架:首先描述 LLM 在六种不同模态下的数据生成应用;更重要的是,从质量与可信度两个维度,系统化分类内在评估指标,实现从依赖下游任务表现的外在评估,转向基于数据自身属性的内在评估。利用该评估体系,分析各模态代表性生成方法的实验结果,揭示当前评估实践中的显著缺陷,并据此向社区提出具体改进建议。最后,框架还给出了合成数据在不同模态中实际应用的方法路径。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have emerged as powerful tools for generating data across various modalities. By transforming data from a scarce resource into a controllable asset, LLMs mitigate the bottlenecks imposed by the acquisition costs of real-world data for model training, evaluation, and system iteration. However, ensuring the high quality of LLM-generated synthetic data remains a critical challenge. Existing research primarily focuses on generation methodologies, with limited direct attention to the quality of the resulting data. Furthermore, most studies are restricted to single modalities, lacking a unified perspective across different data types. To bridge this gap, we propose the \textbf{LLM Data Auditor framework}. In this framework, we first describe how LLMs are utilized to generate data across six distinct modalities. More importantly, we systematically categorize intrinsic metrics for evaluating synthetic data from two dimensions: quality and trustworthiness. This approach shifts the focus from extrinsic evaluation, which relies on downstream task performance, to the inherent properties of the data itself. Using this evaluation system, we analyze the experimental evaluations of representative generation methods for each modality and identify substantial deficiencies in current evaluation practices. Based on these findings, we offer concrete recommendations for the community to improve the evaluation of data generation. Finally, the framework outlines methodologies for the practical application of synthetic data across different modalities.

大模型数据评估可信度合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。