用大模型自动生成有上下文的多维数据故事
MDSF: Context-Aware Multi-Dimensional Data Storytelling Framework based on Large language Model
- 基于微调大模型构建上下文感知的多维度叙事框架
- 在多个数据集上提升洞察排序准确率与故事连贯性
- 适合需要自动化分析报告生成的业务场景
随着数据量激增和大数据技术发展,对高效自动化数据分析与叙事的需求日益增长。现有自动化分析系统在利用大语言模型(LLMs)进行洞察发现、增强分析和数据叙事方面仍面临挑战。本文提出基于大语言模型的多维数据叙事框架(MDSF),实现自动化洞察生成与上下文感知叙事。该框架融合先进预处理技术、增强分析算法及独特评分机制,用于识别和优先排序可操作洞察。通过微调大模型提升上下文理解能力,生成叙述仅需最少人工干预。架构还包含基于代理的机制,支持实时叙事延续控制。实验结果表明,MDSF在多个数据集上优于现有方法,在洞察排序准确性、描述质量与叙事连贯性方面表现更优。评估显示其能自动化复杂分析任务,降低解释偏差,提升用户满意度。用户研究进一步证明其在内容结构优化、结论提取和细节丰富度方面的实用价值。
原文摘要 · Abstract (English)
The exponential growth of data and advancements in big data technologies have created a demand for more efficient and automated approaches to data analysis and storytelling. However, automated data analysis systems still face challenges in leveraging large language models (LLMs) for data insight discovery, augmented analysis, and data storytelling. This paper introduces the Multidimensional Data Storytelling Framework (MDSF) based on large language models for automated insight generation and context-aware storytelling. The framework incorporates advanced preprocessing techniques, augmented analysis algorithms, and a unique scoring mechanism to identify and prioritize actionable insights. The use of fine-tuned LLMs enhances contextual understanding and generates narratives with minimal manual intervention. The architecture also includes an agent-based mechanism for real-time storytelling continuation control. Key findings reveal that MDSF outperforms existing methods across various datasets in terms of insight ranking accuracy, descriptive quality, and narrative coherence. The experimental evaluation demonstrates MDSF's ability to automate complex analytical tasks, reduce interpretive biases, and improve user satisfaction. User studies further underscore its practical utility in enhancing content structure, conclusion extraction, and richness of detail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。