arXiv:2410.12878cs.CLcs.AI2024-10被引 2

评测开源模型在表格转文本中的表现,发现示例能显著提升效果。

Towards More Effective Table-to-Text Generation: Assessing In-Context Learning and Self-Evaluation with Open-Source Models

  • 用上下文示例提升语言模型的表格转文本能力
  • 示例数量越多,生成文本质量越高,但效果饱和明显
  • 自评方法有潜力但与人工评价仍有差距,适合快速评估

表格处理是自然语言处理中的关键任务,近年来受益于语言模型(LMs)的发展。然而,当前开源模型在表格转文本(将结构化数据转化为连贯叙述文本)方面的表现仍需深入研究。本研究在多个基准数据集上评估了不同上下文学习策略的有效性,重点考察提供示例对模型的影响。更重要的是,我们分析了一个真实应用场景,提供了实用洞见。为补充传统评估指标,采用大语言模型(LLM)自评方法,结合思维链推理,评估其与人类对齐指标(如BERTScore)的相关性。结果表明,示例显著提升生成质量,且随着示例增多,性能持续改善,但存在饱和现象;尽管LLM自评具潜力,但其与人类判断的一致性仍有待提高,提示需发展更可靠的评估方法。

原文摘要 · Abstract (English)

Table processing, a key task in natural language processing, has significantly benefited from recent advancements in language models (LMs). However, the capabilities of LMs in table-to-text generation, which transforms structured data into coherent narrative text, require an in-depth investigation, especially with current open-source models. This study explores the effectiveness of various in-context learning strategies in LMs across benchmark datasets, focusing on the impact of providing examples to the model. More importantly, we examine a real-world use case, offering valuable insights into practical applications. To complement traditional evaluation metrics, we employ a large language model (LLM) self-evaluation approach using chain-of-thought reasoning and assess its correlation with human-aligned metrics like BERTScore. Our findings highlight the significant impact of examples in improving table-to-text generation and suggest that, while LLM self-evaluation has potential, its current alignment with human judgment could be enhanced. This points to the need for more reliable evaluation methods.

表格生成自评机制开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。