提出SARAL框架,实现跨语言文档集合检索,超越传统排序列表。
Beyond Ranked Lists: The SARAL Framework for Cross-Lingual Document Set Retrieval
- 聚焦文档集合而非排序列表的跨语言检索方法
- 在六组测试中五组表现优于其他团队,覆盖三种语言
- 适合需要完整信息集合的多语言情报分析场景
机器翻译用于英语检索任意语言信息(MATERIAL)是IARPA推动的跨语言信息检索(CLIR)计划。本文详细描述了信息科学研究所(ISI)在MATERIAL项目中的跨语言摘要与领域自适应检索系统(SARAL)方案。我们提出一种新型方法,旨在检索与查询相关的文档集合,而非仅生成排序文档列表。在MATERIAL第三阶段评估中,SARAL在六组测试条件中的五组表现领先,涵盖波斯语、哈萨克语和格鲁吉亚语三种语言。
原文摘要 · Abstract (English)
Machine Translation for English Retrieval of Information in Any Language (MATERIAL) is an IARPA initiative targeted to advance the state of cross-lingual information retrieval (CLIR). This report provides a detailed description of Information Sciences Institute's (ISI's) Summarization and domain-Adaptive Retrieval Across Language's (SARAL's) effort for MATERIAL. Specifically, we outline our team's novel approach to handle CLIR with emphasis in developing an approach amenable to retrieve a query-relevant document \textit{set}, and not just a ranked document-list. In MATERIAL's Phase-3 evaluations, SARAL exceeded the performance of other teams in five out of six evaluation conditions spanning three different languages (Farsi, Kazakh, and Georgian).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。