arXiv:2601.03496cs.IRcs.CL2026-01

构建航天领域信息检索基准,区分词汇与语义匹配能力。

STELLA: Self-Reflective Terminology-Aware Framework for Building an Aerospace Information Retrieval Benchmark

  • 提出自反思术语感知框架,系统构建航天文档检索数据集。
  • 设计术语一致与术语无关两类查询,分离评估模型词汇与语义能力。
  • 适合关注航天文本检索、嵌入模型评估的研究者使用。

航空航天任务高度依赖技术文档的搜索与复用,但缺乏反映该领域术语特征和查询意图的公开信息检索(IR)基准。本文提出STELLA(Self-Reflective TErminoLogy-Aware Framework for BuiLding an Aerospace Information Retrieval Benchmark)框架,基于NASA技术报告服务器(NTRS)文档,通过文档版面检测、段落切分、术语词典构建、合成查询生成及跨语言扩展的系统化流程,构建了STELLA基准。框架生成两类查询:术语一致查询(TCQ)包含术语原文,用于评估词汇匹配;术语无关查询(TAQ)使用术语描述,用于评估语义匹配。结合链式密度(CoD)与自反思方法提升查询质量,并采用混合跨语言扩展策略模拟真实用户查询行为。在七个嵌入模型上的评估表明,大型解码器模型在语义理解上表现最强,而BM25等词汇匹配方法在强调精确术语匹配的任务中仍具竞争力。STELLA基准为航天领域IR模型的可靠评估与改进提供可复现基础。数据集地址:https://huggingface.co/datasets/telepix/STELLA。

原文摘要 · Abstract (English)

Tasks in the aerospace industry heavily rely on searching and reusing large volumes of technical documents, yet there is no public information retrieval (IR) benchmark that reflects the terminology- and query-intent characteristics of this domain. To address this gap, this paper proposes the STELLA (Self-Reflective TErminoLogy-Aware Framework for BuiLding an Aerospace Information Retrieval Benchmark) framework. Using this framework, we introduce the STELLA benchmark, an aerospace-specific IR evaluation set constructed from NASA Technical Reports Server (NTRS) documents via a systematic pipeline that comprises document layout detection, passage chunking, terminology dictionary construction, synthetic query generation, and cross-lingual extension. The framework generates two types of queries: the Terminology Concordant Query (TCQ), which includes the terminology verbatim to evaluate lexical matching, and the Terminology Agnostic Query (TAQ), which utilizes the terminology's description to assess semantic matching. This enables a disentangled evaluation of the lexical and semantic matching capabilities of embedding models. In addition, we combine Chain-of-Density (CoD) and the Self-Reflection method with query generation to improve quality and implement a hybrid cross-lingual extension that reflects real user querying practices. Evaluation of seven embedding models on the STELLA benchmark shows that large decoder-based embedding models exhibit the strongest semantic understanding, while lexical matching methods such as BM25 remain highly competitive in domains where exact lexical matching technical term is crucial. The STELLA benchmark provides a reproducible foundation for reliable performance evaluation and improvement of embedding models in aerospace-domain IR tasks. The STELLA benchmark can be found in https://huggingface.co/datasets/telepix/STELLA.

信息检索航天文本嵌入模型术语匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。