构建真实金融数据叙述基准,评估模型讲好数据故事的能力。
DataTales: A Benchmark for Real-World Intelligent Data Narration
- 用4900份财报与市场数据配对,模拟真实场景
- 模型在专业术语理解与深度分析上表现不足
- 适合研究数据叙事、金融AI的开发者和评测者
我们提出DataTales,一个用于评估语言模型在数据叙述任务中表现的新基准。该任务对将复杂表格数据转化为可读叙述至关重要。现有基准难以反映实际应用所需的分析复杂性。DataTales通过提供4.9k份财务报告及其对应的市场数据,展示了模型需具备清晰叙述能力、处理大规模数据集,并理解领域内专业术语。实验表明,当前语言模型在实现足够精确性和分析深度方面仍面临重大挑战,为未来模型开发与评估方法指明了方向。
原文摘要 · Abstract (English)
We introduce DataTales, a novel benchmark designed to assess the proficiency of language models in data narration, a task crucial for transforming complex tabular data into accessible narratives. Existing benchmarks often fall short in capturing the requisite analytical complexity for practical applications. DataTales addresses this gap by offering 4.9k financial reports paired with corresponding market data, showcasing the demand for models to create clear narratives and analyze large datasets while understanding specialized terminology in the field. Our findings highlights the significant challenge that language models face in achieving the necessary precision and analytical depth for proficient data narration, suggesting promising avenues for future model development and evaluation methodologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。