arXiv:2609.07629cs.IRcs.AI2026-09

梳理表格洞察提取现状,提出统一框架OpenTI并指明未来方向。

Open Tabular Insight Extraction: Where Do We Stand, and Where Should We Go?

论文配图:Open Tabular Insight Extraction: Where Do We Stand, and Where Should We Go?
图 1 · 摘自论文原文
  • 构建统一框架OpenTI,整合多领域知识以实现端到端表格洞察提取。
  • 发现现有系统仅覆盖分析环节,且基准测试不适用于开放场景。
  • 适合关注数据智能、人机交互与数据库融合研究的学者参考。

将大规模表格数据湖中的知识民主化获取正成为核心研究挑战。该领域的研究不断推进并拓展范围,逐步提供满足用户洞察需求的端到端组件。然而,这些工作仍分散在不同社区中,各自采用不同的问题定义范式,如表格问答、文本转SQL和数据分析代理,跨任务标签的引用频率仅为同标签内的六分之一。为弥合这一分歧,本文提出一个整体性框架——开放表格洞察提取(Open Tabular Insight Extraction, OpenTI),从用户所需的分析知识、从表格语料库中推导知识的流程,以及结果对用户的适用性三个基本层面进行形式化。该框架整合信息检索、自然语言处理、机器学习、数据库与人机交互等领域的理论与术语,并在此基础上对致力于OpenTI的系统与基准进行系统性回顾与分析。结果显示,当前系统未覆盖完整的端到端范围,主要聚焦于分析本身;而基准测试大多不适合开放环境评估,因输入预设了表格知识,且验证机制与实际场景不符。最后,本文提炼出面向OpenTI系统的研发、评估与交互范式的未来研究议程。论文配套交互式网页见:https://open-tabular-insight-extraction.github.io。

原文摘要 · Abstract (English)

Democratizing access to the knowledge held in large corpora of tables such as data lakes is emerging as a central research challenge. Research in this space is advancing and broadening in scope, increasingly supplying the components to satisfy a person's insight need end-to-end. Yet these efforts remain fragmented across communities that frame the problem under their own conventions, such as table question answering, text-to-SQL, and data analysis agents, with works six times as likely to cite within the same task label as across labels. To bring these communities onto common ground, we establish a holistic framework for this pursuit, which we refer to as Open Tabular Insight Extraction (OpenTI). We formalize OpenTI from first principles around the analytical knowledge a person needs, the procedure for deriving it from a corpus of tables, and how well a result serves the person who sought it. In doing so we consolidate frameworks and terminology across information retrieval, natural language processing, machine learning, databases, and human-computer interaction, and apply this grounding in a systematic review and analysis of systems and benchmarks that work towards OpenTI. We find that current systems do not cover the end-to-end scope of OpenTI, mainly focusing on the analysis itself, and that benchmarks are largely unfit for evaluations in an open setting as inputs presuppose knowledge of tables, and validation mechanisms do not match the setup. Finally, we distill a research agenda towards OpenTI systems, evaluation, and interaction paradigms that surface the insights users need. An interactive companion to our paper is available at https://open-tabular-insight-extraction.github.io.

表格洞察数据智能人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。