评估持续学习方法时,时间与适配容量会影响结果排名,需看完整表现轨迹。
Beyond Endpoint Scores: Time- and Capacity-Conditioned Evaluation of Continual Knowledge Updating

- 对比周期性结构与累计回放,变化评估时间与适配容量
- 同一方法在不同条件下排名反转,最高差达11.6分
- 建议报告性能轨迹,避免片面结论,适合方法对比研究者
持续知识更新方法常仅依据单一检查点和固定适配器秩判断优劣,但此方式可能不足。固定周期结构,对比其在24个月Wikidata流上与累计回放的表现,同时改变评估月份、回放LoRA秩及查询形式。结果显示,方法排名随条件变化:在Qwen2.5-1.5B上,周期结构对秩8回放有5.0分优势,却在秩72时转为11.6分劣势;高秩下整合对齐终点显示平局,而时间平均回放领先9-13分。相同现象出现在Llama-3.2-1B及保留的改写样本中。表明持续更新方法的排名受评估时机与回放侧适应能力共同影响。因此建议报告性能轨迹与容量扫描结果,仅当排序稳定时才宣布优胜者;否则应说明优胜区域与保留-稳定性-成本前沿。依此协议,周期结构仅为低更新成本方案,非质量优胜者。
原文摘要 · Abstract (English)
Continual knowledge-updating methods are often declared superior from one final checkpoint and one conventional adapter rank. We show that this can be insufficient to identify the better operating point. Holding a periodic hierarchy fixed, we compare it with cumulative replay over a 24-month Wikidata stream while varying evaluation month, replay LoRA rank, and query formulation. The apparent winner changes across this region: on Qwen2.5-1.5B, the hierarchy's 5.0-point advantage over rank-8 replay becomes an 11.6-point deficit against rank-72 replay, and at high ranks a consolidation-aligned endpoint can suggest a tie while time-averaged replay leads by 9-13 points. The same rank-conditioned reversal appears on Llama-3.2-1B and held-out paraphrases. These results show that method ranking in continual updating can depend jointly on when performance is measured and how much replay-side adaptation capacity the baseline receives. We therefore propose reporting trajectories and capacity sweeps, and declaring a robust winner only when the ordering is stable across the evaluation region; otherwise, comparisons should report winner regions and retention-stability-cost frontiers. Under this protocol, the periodic hierarchy is a lower-update-cost operating point, not a quality winner.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。