arXiv:2507.05865cs.IRcs.DB2025-07

动态高维数据下,学习型索引如何高效扩容并超越静态版本。

On the Costs and Benefits of Learned Indexing for Dynamic High-Dimensional Data: Extended Version

  • 通过节点分裂与扩展实现静态学习索引的动态化
  • 动态索引随数据增长,总成本更低且性能更优
  • 适合处理持续增长的高维数据库场景

学习型索引在动态扩展数据集时缺乏适应性,本文提出通过节点分裂与拓宽等操作,将静态学习索引扩展为可动态适应新数据的结构。同时引入摊销成本模型,综合评估查询性能与索引构建成本,实验表明:随着数据库规模增长,动态学习索引的总体成本迅速低于静态版本。该文为DAWAK 2025会议论文的扩展版。

原文摘要 · Abstract (English)

One of the main challenges within the growing research area of learned indexing is the lack of adaptability to dynamically expanding datasets. This paper explores the dynamization of a static learned index for complex data through operations such as node splitting and broadening, enabling efficient adaptation to new data. Furthermore, we evaluate the trade-offs between static and dynamic approaches by introducing an amortized cost model to assess query performance in tandem with the build costs of the index structure, enabling experimental determination of when a dynamic learned index outperforms its static counterpart. We apply the dynamization method to a static learned index and demonstrate that its superior scaling quickly surpasses the static implementation in terms of overall costs as the database grows. This is an extended version of the paper presented at DAWAK 2025.

学习型索引动态数据高维数据数据库优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。