输入格式影响大模型体育赛事摘要的准确性,结构化数据可大幅减少事实错误。
Input Matters: Evaluating Input Structure's Impact on LLM Summaries of Sports Play-by-Play
- 对比三种输入格式:行结构、JSON和无结构,评估其对大模型摘要的影响。
- 使用JSON输入使错误率降低65%~69%,优于其他两种格式。
- 研究适合关注大模型输出可信度的体育内容生成与部署人员。
在体育报道等对准确性要求高的领域部署大模型时,生成文本可能无法忠实反映输入数据,成为一大担忧。本文量化分析了输入结构对大模型生成NBA比赛实况摘要中幻觉及其他事实错误的影响,涵盖三种格式:行结构、JSON和无结构。研究手动标注了180场游戏摘要中的3,312个事实错误,涉及Llama-3.1-70B和Qwen2.5-72B两个模型。结果显示,输入结构影响显著:相较于无结构输入,使用JSON输入使Llama错误率降低69%,Qwen降低65%;行结构输入分别降低54%和51%。双因素重复测量方差分析表明,输入结构解释了超过80%的错误率变异,且所有输入格式间差异均具有统计学意义。
原文摘要 · Abstract (English)
A major concern when deploying LLMs in accuracy-critical domains such as sports reporting is that the generated text may not faithfully reflect the input data. We quantify how input structure affects hallucinations and other factual errors in LLM-generated summaries of NBA play-by-play data, across three formats: row-structured, JSON and unstructured. We manually annotated 3,312 factual errors across 180 game summaries produced by two models, Llama-3.1-70B and Qwen2.5-72B. Input structure has a strong effect: JSON input reduces error rates by 69% for Llama and 65% for Qwen compared to unstructured input, while row-structured input reduces errors by 54% for Llama and 51% for Qwen. A two-way repeated measures ANOVA shows that input structure accounts for over 80% of the variance in error rates, with Tukey HSD post hoc tests confirming statistically significant differences between all input formats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。