用新基准测试大模型在体育长文中的表格总结能力,发现多主体记忆是核心瓶颈
Moneyball with LLMs: Analyzing Tabular Summarization in Sports Narratives
- 构建体育领域长文本表格总结诊断基准SporTabSet
- 分解策略提升准确率但主要靠减少主体干扰而非算术增强
- 模型易受表面线索误导,常出现幻觉和角色混淆
大型语言模型(LLM)在表格总结任务中依赖大量提示工程、分解流程或实体级中间表示以获得良好表现。尽管有效,这些方法计算成本高,且难以揭示模型在长期演化叙事中维持状态的能力。本文提出SporTabSet,一个针对两个互补体育领域的长上下文表格总结诊断基准,需跟踪多个主体并按领域规则聚合统计数据。利用SporTabSet,系统评估了多种长上下文LLM上的分解策略。结果表明,尽管分解显著提升准确率和数值保真度,但增益主要源于缓解多主体干扰,而非局部算术能力的改善。鲁棒性实验进一步揭示模型对表面线索高度敏感,存在结构化失败模式,包括幻觉、遗漏和角色混淆。这些发现一致指出多主体记忆是长上下文表格生成的关键瓶颈,强调诊断评估应成为可扩展、高效且可靠的表格总结模型研发的先决条件。
原文摘要 · Abstract (English)
Large language model (LLM) approaches to tabular summarization rely on extensive prompt engineering, decomposition pipelines, or entity-level intermediate representations to achieve strong performance. While effective, these strategies are computationally expensive and offer limited insight into how well models maintain state over long, evolving narratives. We introduce SPORTABSET, a diagnostic benchmark for long-context tabular summarization across two complementary sports domains that require tracking multiple entities and aggregating statistics under domain-specific rules. Using SporTabSet, we systematically evaluate decomposition-based strategies across several long context LLMs. Results show that although decomposition substantially improves accuracy and numerical fidelity, gains stem mainly from dissecting multi-entity interference rather than improved local arithmetic. Robustness experiments further reveal high sensitivity to surface-level cues with structured failures, including hallucination, omission, and role confusion. Together, these findings identify consistent multientity memory as a key bottleneck in long context table generation, motivating diagnostic evaluation as a prerequisite for scalable, efficient and reliable tabular summarization models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。