不依赖关键词匹配,实现跨语言跨模态记忆精准检索
Wontopos Tablet 2: Measuring Multilingual and Multimodal Memory Retrieval Without Lexical Matching

- 纯向量检索路径无词法匹配与语言模型
- 跨语言图文检索准确率达95.2%(对比BM25仅19.0%)
- 适合多语言、无文本描述场景下的长时记忆系统研究
我们评估了Tablet-2——一个用于语言模型的生产级长期记忆引擎,在现有文本基准和无文字标注图片的跨语言检索任务上表现。其检索路径不包含词法匹配、关键词评分或自有的语言模型。在LongMemEval-S(500个问题)上得分为95.7% [93.4, 97.1];在BEAM-1M(700个问题,221万条存储记忆)上为67.5% [64.8, 70.2]。这些是问题采样区间,而非运行间波动,后者窄一个数量级。固定引擎、语料、设置与评测者,仅改变阅读器可使LongMemEval-S波动2.0分,仅改变重问预算则使BEAM-1M波动8.9分,远超多数已有报告差距,因此仅作定位参考。多模态实验中,相较配置至极限的BM25,我们在70个语言单元中平均召回率@5达95.2%,而BM25仅为19.0%,且在54个单元中为零。对无字幕照片,词法方法无文档可评分。开放密集基线在300张跨语言跨模态图像(Crossmodal-3600,14语言)上显示:密度未带来语言独立性——一模型英语91.0%、俄语仅4.7%;多语言变体在泰卢固语和斯瓦希里语上崩溃。本研究语言间波动14.0,对比分别为27.5和27.7。三项反向结果被同等报告:低资源语言显著下降(斯瓦希里53.0%,泰卢固64.0%),附加字幕使跨语言检索下降11.4分,一项设置遗漏导致韩语top-1下降37分但其余九种语言不变。
原文摘要 · Abstract (English)
We measure tablet-2, a production long-term memory engine for language models, on the text benchmarks the field already uses and on cross-lingual retrieval of photographs stored with no text at all. Its retrieval path contains no lexical matching, no keyword scoring, and no language model of its own. On LongMemEval-S (500 questions) it scores 95.7% [93.4, 97.1]; on BEAM-1M (700 questions, 2.21M stored memories) 67.5% [64.8, 70.2]. Those are question-sampling intervals, not the run-to-run spread, which is an order of magnitude narrower. Most of the paper is about how little they mean alone. Holding engine, corpus, settings and judge fixed, changing only the reader moves LongMemEval-S by 2.0 points; changing only the re-ask budget moves BEAM-1M by 8.9. Neither is stated in the reports we compare against, and the second exceeds most gaps there, so we give that table as a placement and not a ranking. For the multimodal axis we run two controls. Against BM25, configured as strongly as we could, we reach 95.2% mean recall@5 over 70 store-and-query language cells where BM25 reaches 19.0% and is exactly zero in 54. On captionless photographs a lexical method has no document to score at all. Open dense baselines on 300 Crossmodal-3600 photographs in 14 languages show that density confers no language independence: one scores 91.0% on English and 4.7% on Russian from identical image vectors, and a multilingual variant collapses on Telugu and Swahili. Our spread across languages is 14.0 against their 27.5 and 27.7. Three results run against us and are reported at equal weight: low-resource languages degrade sharply (Swahili 53.0%, Telugu 64.0%), attaching captions lowers cross-lingual retrieval by 11.4 points, and one setting omitted into one stage of our own retrieval cost 37 points of Korean top-1 accuracy while leaving nine languages untouched.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。