arXiv:2510.13598cs.CL2025-10被引 1

FreshTab通过实时生成新数据,解决大模型评估中的数据污染问题。

FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation

  • 从维基百科实时生成表到文本的评测数据,避免模型训练数据污染。
  • 新数据生成的模型输出在自动指标上表现更差,但人评和模型评无显著差异。
  • 支持多语言动态采集,可实现领域平衡的公平评估,适合评估模型泛化能力。

表到文本生成(从表格中提炼见解)是一项需要精确数据分析的挑战性任务。现有基准评估受大型语言模型(LLM)训练数据污染及领域不平衡影响。我们提出FreshTab,一种基于维基百科实时生成表到文本评测数据的方法,以应对LLM数据污染问题并实现领域敏感评估。尽管非英语表到文本数据集稀缺,FreshTab可按需收集多语言数据(本文实验包含德语、俄语和法语)。我们发现,使用本方法获取的近期表格生成的见解在自动指标上表现明显更差,但未体现在人类与模型评价中。所有评估均显示出领域效应,表明领域均衡的基准更具挑战性。

原文摘要 · Abstract (English)

Table-to-text generation (insight generation from tables) is a challenging task that requires precision in analyzing the data. In addition, the evaluation of existing benchmarks is affected by contamination of Large Language Model (LLM) training data as well as domain imbalance. We introduce FreshTab, an on-the-fly table-to-text benchmark generation from Wikipedia, to combat the LLM data contamination problem and enable domain-sensitive evaluation. While non-English table-to-text datasets are limited, FreshTab collects datasets in different languages on demand (we experiment with German, Russian and French in addition to English). We find that insights generated by LLMs from recent tables collected by our method appear clearly worse by automatic metrics, but this does not translate into LLM and human evaluations. Domain effects are visible in all evaluations, showing that a~domain-balanced benchmark is more challenging.

表到文本数据污染多语言评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。