提升科学文献长时检索效果,关键在全文+时间维度融合。
Submitted and Diagnostic Analysis of Full-Text Temporal Retrieval for LongEval-Sci

- 用全文匹配与时间信息融合,增强长期检索稳定性。
- 在三个时间快照上均达最优nDCG@10(最高0.285)。
- 适合关注系统长期可用性的信息检索研究者。
LongEval-Sci评估在文献集合随时间演化的场景下,检索系统对当前语料的有效性及长期可用性。本文报告了LongEval-Sci 2026官方任务1的结果与开发诊断。对比了官方PyTerrier BM25和Qwen3密集基线,以及全文字典BM25、加性与路由变体、时间感知全文字检索、时间+引用检索、RM3查询扩展、交叉编码重排和倒数排名融合(RRF)。在官方DCTR评估中,时间化全文字检索表现最佳:FT BM25+temporal与FT BM25+temporal+citation在三个快照上均取得最高ARP,nDCG@10分别为0.285、0.267、0.180;将快照3相对变化从基线的0.481降至0.368。引用特征表现与仅时间变体相当,未带来显著提升。内部快照1诊断显示,全文字BM25为最强单模型(DCTR nDCG@10=0.3302,MAP=0.2853),RRF实现最佳深层召回(Recall@1000=0.9667),部分未校准叠加会严重损害前排质量。因此结论为:全文字检索是强基础,时间整合可提升长期效果,引用证据仍需更清晰的消融与校准。此外,还基于数据摄入速度与陈旧覆盖漂移进行了定性每周更新监控分析。
原文摘要 · Abstract (English)
LongEval-Sci evaluates scientific retrieval under collection change, where a system should be effective on the current corpus and remain usable as documents accumulate over time. This paper reports both official Task 1 results and development diagnostics for LongEval-Sci 2026. We compare the official PyTerrier BM25 and Qwen3 dense baselines with full-text BM25, additive and router variants, temporal full-text retrieval, temporal+citation retrieval, RM3 query expansion, cross-encoder reranking, and reciprocal rank fusion (RRF). In the official DCTR evaluation, the temporalized full-text runs are our strongest submissions: FT BM25+temporal and FT BM25+temporal+citation obtain the best ARP on all three snapshots (0.285, 0.267, and 0.180 nDCG@10) and reduce snapshot-3 relative change from 0.481 for the BM25 pivot to 0.368. Citation features match the temporal-only variant but do not provide a measurable additional gain in the official summary. Our internal snapshot-1 diagnostics show a complementary pattern: full-text BM25 is the strongest single development retriever (DCTR nDCG@10 = 0.3302, MAP = 0.2853), RRF gives the best deep recall (Recall@1000 = 0.9667), and some uncalibrated overlays can sharply degrade top-rank quality. We therefore conclude that full-text retrieval is the strongest foundation, temporal integration can improve official longitudinal effectiveness when applied to that foundation, and citation evidence still requires cleaner ablation and calibration. Beyond ranking, we also report a qualitative weekly IR-system update-monitoring analysis based on ingestion velocity and stale-coverage drift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。