首个面向实时信息推荐的基准数据集,解决何时何地给对信息的问题。
JIR-Arena: The First Benchmark Dataset for Just-in-time Information Recommendation
- 通过多源输入融合与静态知识库评估,构建客观的推荐评测框架。
- 现有模型虽能合理模拟用户需求,但召回率与内容检索效率仍不足。
- 适合研究智能可穿戴设备、实时推荐系统的学者与开发者。
即时信息推荐(JIR)旨在用户需要时精准推送最相关的信息,填补知识空白,提升日常决策效率。随着基础模型轻量化部署和智能可穿戴设备普及,持续在线的JIR助手成为可能。然而,当前缺乏对JIR任务的明确定义与评估体系。为此,我们首次提出JIR任务的数学定义及配套评估指标,并发布首个多模态基准数据集JIR-Arena。该数据集涵盖多样化的高信息需求场景,用于评估系统在三方面表现:准确推断用户信息需求、及时提供相关推荐、避免干扰性内容。为应对用户需求主观性与系统变量不可控问题,JIR-Arena结合多人与大模型输入以近似信息需求分布,利用静态知识库快照评估推荐质量,并采用多轮多实体验证框架提升客观性与泛化能力。我们还实现了一个可处理实时信息流的基线系统。在JIR-Arena上的评估表明,基于大模型的系统虽能较准确模拟用户需求,但在召回率和有效内容检索上仍有挑战。为推动该领域发展,代码与数据已完全开源。
原文摘要 · Abstract (English)
Just-in-time Information Recommendation (JIR) is a service designed to deliver the most relevant information precisely when users need it, , addressing their knowledge gaps with minimal effort and boosting decision-making and efficiency in daily life. Advances in device-efficient deployment of foundation models and the growing use of intelligent wearable devices have made always-on JIR assistants feasible. However, there has been no systematic effort to formally define JIR tasks or establish evaluation frameworks. To bridge this gap, we present the first mathematical definition of JIR tasks and associated evaluation metrics. Additionally, we introduce JIR-Arena, a multimodal benchmark dataset featuring diverse, information-request-intensive scenarios to evaluate JIR systems across critical dimensions: i) accurately inferring user information needs, ii) delivering timely and relevant recommendations, and iii) avoiding irrelevant content that may distract users. Developing a JIR benchmark dataset poses challenges due to subjectivity in estimating user information needs and uncontrollable system variables affecting reproducibility. To address these, JIR-Arena: i) combines input from multiple humans and large AI models to approximate information need distributions; ii) assesses JIR quality through information retrieval outcomes using static knowledge base snapshots; and iii) employs a multi-turn, multi-entity validation framework to improve objectivity and generality. Furthermore, we implement a baseline JIR system capable of processing real-time information streams aligned with user inputs. Our evaluation of this baseline system on JIR-Arena indicates that while foundation model-based JIR systems simulate user needs with reasonable precision, they face challenges in recall and effective content retrieval. To support future research in this new area, we fully release our code and data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。