arXiv:2503.15509cs.HCcs.CL2025-03被引 1

将数据转化为自然叙事,让大模型更准确地表达数值信息

Representing data in words: A context engineering approach

  • 通过上下文工程生成符合语境的描述性文本,提升模型对数值的理解
  • 在足球球员评估、人格测试等场景中实现高准确率的数据转述
  • 适合需要可解释数据报告的研究者与决策者使用

大型语言模型在诸多应用中展现出巨大潜力,但如何生成忠实反映数据的可靠文本仍具挑战。尽管先前研究证明任务特定的上下文学习和知识增强能提升性能,大模型在解析与推理数值数据方面依然表现不足。为此,本文提出wordalisations方法,将数据洞察抽象为风格自然的叙述性文本,如同可视化使数字易于理解。我们将其应用于足球球员选拔、人格测试和国际调查数据三个场景。由于该任务缺乏标准基准,我们采用大模型作为评判者和人工评判进行评估,结果表明wordalisations生成的文本既生动又准确。同时,我们还提出了开放透明的数据传播实践建议。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable potential across a broad range of applications. However, producing reliable text that faithfully represents data remains a challenge. While prior work has shown that task-specific conditioning through in-context learning and knowledge augmentation can improve performance, LLMs continue to struggle with interpreting and reasoning about numerical data. To address this, we introduce wordalisations, a methodology for generating stylistically natural narratives from data. Much like how visualisations display numerical data in a way that is easy to digest, wordalisations abstract data insights into descriptive texts. To illustrate the method's versatility, we apply it to three application areas: scouting football players, personality tests, and international survey data. Due to the absence of standardized benchmarks for this specific task, we conduct LLM-as-a-judge and human-as-a-judge evaluations to assess accuracy across the three applications. We found that wordalisation produces engaging texts that accurately represent the data. We further describe best practice methods for open and transparent development of communication about data.

数据叙事大模型自然语言生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。