针对表格不同字段设计混合匹配策略,提升表格检索效果
Tailoring Table Retrieval from a Field-aware Hybrid Matching Perspective
- 按字段区分匹配粒度,细胞用词级匹配,标题用句级匹配
- 在NQ-TABLES和OTT-QA上超越现有最佳方法
- 适合需要精准表格信息检索的研究者与工程师
表格检索对于通过结构化数据获取信息至关重要,但相较于文本检索研究较少。表格的行/列结构及各字段(包括标题、表头、单元格)具有独特性,不同字段存在差异化的匹配偏好:单元格因碎片化和细节丰富,更适合词/短语级细粒度匹配,而标题则更适配句子/段落级粗粒度匹配。因此,需设计面向表格的检索器以满足各字段的匹配需求。为此,我们提出一种场域感知的混合匹配检索器——THYME(Table-tailored HYbrid Matching rEtriever),从场域感知的混合匹配视角解决表格检索问题。在NQ-TABLES和OTT-QA两个基准测试上的实证结果表明,THYME显著优于现有最先进基线模型。全面分析验证了不同字段间的匹配偏好差异,并支持了THYME的设计合理性。
原文摘要 · Abstract (English)
Table retrieval, essential for accessing information through tabular data, is less explored compared to text retrieval. The row/column structure and distinct fields of tables (including titles, headers, and cells) present unique challenges. For example, different table fields have varying matching preferences: cells may favor finer-grained (word/phrase level) matching over broader (sentence/passage level) matching due to their fragmented and detailed nature, unlike titles. This necessitates a table-specific retriever to accommodate the various matching needs of each table field. Therefore, we introduce a Table-tailored HYbrid Matching rEtriever (THYME), which approaches table retrieval from a field-aware hybrid matching perspective. Empirical results on two table retrieval benchmarks, NQ-TABLES and OTT-QA, show that THYME significantly outperforms state-of-the-art baselines. Comprehensive analyses confirm the differing matching preferences across table fields and validate the design of THYME.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。