arXiv:2505.10043cs.IRcs.AI2025-05

用合成语义信息提升文本查图准确率,解决真实商业场景下理解难题。

Boosting Text-to-Chart Retrieval through Training with Synthesized Semantic Insights

  • 构建多层次语义洞察生成流水线,自动提炼图表的视觉、统计与应用层信息
  • 在新基准CRBench上,精确查询NDCG@10达66.9%,较最优方法提升11.58%
  • 适合关注图表理解、智能分析系统与数据可视化研究者使用

文本到图表检索使用户可通过自然语言查询查找相关图表,日益受到关注。然而,在真实商业智能(BI)场景中评估模型仍具挑战性,因现有基准无法模拟真实用户查询或测试静态图表图像下的深层语义理解能力。为填补此空白,我们提出CRBench,首个源自真实业务场景的基准,包含21,862张图表和326个查询,采用目标-干扰物范式评估高度相似候选间的判别性检索。测试表明,依赖视觉特征的现有方法表现不佳,难以捕捉图表丰富的分析语义。为此,我们设计了一条语义洞察合成流水线,自动为图表生成三个层次的洞察:视觉模式、统计属性与实际应用。基于该流水线,我们为69,166张图表生成了207,498条语义洞察作为训练数据。通过多层级语义监督,将自然语言查询意图与隐式视觉表征对齐,开发出专用模型ChartFinder,实现深度跨模态推理。实验结果表明,ChartFinder在CRBench上显著优于现有方法:精确查询的NDCG@10最高达66.9%(提升11.58%),模糊查询各项指标平均提升5%。本工作为社区提供了一个亟需的真实评估基准,并展示了增强图表语义理解的强大数据合成范式。

原文摘要 · Abstract (English)

Text-to-chart retrieval, enabling users to find relevant charts via natural language queries, has gained significant attention. However, evaluating models in real-world business intelligence (BI) scenarios is challenging, as current benchmarks fail to simulate realistic user queries or test for deep semantic understanding with static chart images.To address this gap, we introduce CRBench, the first real-world BI-sourced benchmark comprising 21,862 charts and 326 queries, utilizing a Target-and-Distractor paradigm to evaluate discriminative retrieval among highly similar candidates. Testing on CRBench reveals that existing methods, which rely primarily on visual features, perform poorly and fail to capture the rich analytical semantics of charts. To address this performance bottleneck, we propose a semantic insights synthesis pipeline that automatically generates three hierarchical levels of insights for charts: visual patterns, statistical properties, and practical applications. Using this pipeline, we produced 207,498 semantic insights for 69,166 charts as training data. By leveraging this data to bridge the gap between natural language query intent and latent visual representations via multi-level semantic supervision, we develop ChartFinder, a specialized model capable of deep cross-model reasoning. Experimental results show ChartFinder significantly outperforms state-of-the-art methods on CRBench, achieving up to 66.9% NDCG@10 for precise queries (an 11.58% improvement) and an average increase of 5% across nearly all metrics for fuzzy queries. This work provides the community with a much-needed benchmark for realistic evaluation and demonstrates a powerful data synthesis paradigm for enhancing a model's semantic understanding of charts.

文本查图语义理解数据合成图表检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。