arXiv:2608.06614cs.CLcs.AI2026-08

针对间接证据的分类检索难题,提出分因子假设搜索方法提升准确率。

Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval

  • 构建多维语义假设并行搜索,支持结构化查询与验证
  • 在金融和临床编码任务中召回率与排序指标均最优
  • 分因子设计比自由文本集成更有效,且无需逐步优化

大类别检索通常假设输入已明确表达目标概念,但在许多场景中,输入仅为间接证据(如表格单元格),其含义依赖于行、列、数据类型和上下文。这种语义不明确导致检索能力下降,称为‘检索就绪差距’。分析表明,当前索引在语义明确时表现可靠,但对原始证据则将其排在深层。本文提出分因子假设搜索(FHS),在多个命名语义维度上维护部分解释。这些假设支持结构化查询生成、多假设检索与维度级候选验证。在金融分类标注与CodiEsp临床编码任务中,FHS在非人工标注方法中实现了最高的Recall@1、MRR与最终准确率。将分因子假设路径替换为自由文本集成导致头部排名性能最大下降,而序列精炼未能带来优于FHS并行初轮的优势。

原文摘要 · Abstract (English)

Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirect evidence, such as a table cell whose meaning depends on its row, column, datatype, and context. We call this mismatch the retrieval readiness gap. Our analysis shows that the current index retrieves the target reliably when its semantics are explicit, while raw evidence often leaves it deep in the ranking. We propose Factorized Hypothesis Search (FHS), which maintains multiple partial interpretations over named semantic dimensions. These hypotheses support structured query rendering, multi-hypothesis retrieval, and dimension-level candidate verification. On both financial taxonomy tagging and CodiEsp clinical coding tasks, FHS achieves the best Recall@1, MRR, and final accuracy among the non-oracle methods. Replacing the factorized hypothesis path with a free-text ensemble causes the largest drop in head-ranking performance, while sequential refinement provides no additional gain over FHS's strong parallel first round.

检索增强多维推理知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。