让检索器学会识别时间信息,提升时序问答效果。
Efficient Temporal-aware Matryoshka Adaptation for Temporal Information Retrieval
- 用嵌套结构的时序嵌入增强文本表示,保留语义同时加入时间感知。
- 在多个数据集上达到与顶尖方法相当的时序检索效果。
- 可灵活调节精度与效率,适合部署于资源受限场景。
在时序检索增强生成(Temporal RAG)系统中,检索器是关键瓶颈:若无法检索到相关时间上下文,即使大模型推理能力再强,下游生成也会失效。本文提出时序感知嵌套表征学习(TMRL),通过嵌套结构的Matryoshka嵌入引入时序子空间,在增强时序编码的同时保留通用语义表征。实验表明,TMRL能高效适配多种文本嵌入模型,在时序检索与时序RAG任务中表现优于以往基于Matryoshka的非时序方法及主流时序方法,且支持灵活的精度-效率权衡。
原文摘要 · Abstract (English)
Retrievers are a key bottleneck in Temporal Retrieval-Augmented Generation (RAG) systems: failing to retrieve temporally relevant context can degrade downstream generation, regardless of LLM reasoning. We propose Temporal-aware Matryoshka Representation Learning (TMRL), an efficient method that equips retrievers with temporal-aware Matryoshka embeddings. TMRL leverages the nested structure of Matryoshka embeddings to introduce a temporal subspace, enhancing temporal encoding while preserving general semantic representations. Experiments show that TMRL efficiently adapts diverse text embedding models, achieving competitive temporal retrieval and temporal RAG performance compared to prior Matryoshka-based non-temporal methods and prior temporal methods, while enabling flexible accuracy-efficiency trade-offs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。