arXiv:2510.23854cs.CLcs.AI2025-10EMNLP被引 8

评估大模型生成表格自然语言描述的可靠性,提出高效新方法

Can LLMs Narrate Tabular Data? An Evaluation Framework for Natural Language Representations of Text-to-SQL System Outputs

  • 设计组合式评估框架Combo-Eval,融合多种评价方式
  • 减少25%-61%的大模型调用次数,保持高评估精度
  • 构建首个专用数据集NLR-BIRD,适合评测表格转文本任务

在现代工业系统如多轮对话代理中,文本转SQL技术连接自然语言问题与数据库查询。将表格型数据库结果转化为自然语言表述(NLRs)可实现基于聊天的交互。目前,该过程通常由大语言模型(LLMs)完成,但表格结果在自然语言中呈现时的信息丢失或错误尚未被充分研究。本文提出一种新型评估方法Combo-Eval,结合多种现有方法的优势,在提升评估保真度的同时,使大模型调用次数减少25%-61%。配套提出首个针对NLR基准测试的数据集NLR-BIRD。通过人工评估,验证了Combo-Eval在有无真实参考情况下均与人类判断高度一致。

原文摘要 · Abstract (English)

In modern industry systems like multi-turn chat agents, Text-to-SQL technology bridges natural language (NL) questions and database (DB) querying. The conversion of tabular DB results into NL representations (NLRs) enables the chat-based interaction. Currently, NLR generation is typically handled by large language models (LLMs), but information loss or errors in presenting tabular results in NL remains largely unexplored. This paper introduces a novel evaluation method - Combo-Eval - for judgment of LLM-generated NLRs that combines the benefits of multiple existing methods, optimizing evaluation fidelity and achieving a significant reduction in LLM calls by 25-61%. Accompanying our method is NLR-BIRD, the first dedicated dataset for NLR benchmarking. Through human evaluations, we demonstrate the superior alignment of Combo-Eval with human judgments, applicable across scenarios with and without ground truth references.

自然语言生成大模型评估文本转SQL

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。