用不确定性量化提升学习型索引收益估计的可靠性
Can Uncertainty Quantification Improve Learned Index Benefit Estimation?
- 融合自编码器与蒙特卡洛丢弃,联合量化预测不确定性
- 在16个模型中多数表现优于现有方法,6个数据集上消除最差情况
- 适合数据库优化、机器学习部署等需可靠决策的场景
索引调优对数据库性能优化至关重要,核心在于准确高效的收益估计。传统基于what-if工具的方法常因效率低和不准确而受限;学习型模型虽具潜力,却存在不稳定、不可解释及管理复杂等问题。为此,本文提出Beauty——首个具备不确定性感知能力的框架,通过量化学习模型输出的不确定性,并以what-if工具作为补充机制,兼顾可靠性与易用性。创新性地结合自编码器与蒙特卡洛丢弃,针对收益估计任务特性设计联合不确定性量化方法。在16个模型上的实验表明,该方法在多数情况下超越现有技术。六组数据集的索引调优测试显示,应用Beauty后彻底消除最差情况,最佳情况发生次数提升逾三倍。
原文摘要 · Abstract (English)
Index tuning is crucial for optimizing database performance by selecting optimal indexes based on workload. The key to this process lies in an accurate and efficient benefit estimator. Traditional methods relying on what-if tools often suffer from inefficiency and inaccuracy. In contrast, learning-based models provide a promising alternative but face challenges such as instability, lack of interpretability, and complex management. To overcome these limitations, we adopt a novel approach: quantifying the uncertainty in learning-based models' results, thereby combining the strengths of both traditional and learning-based methods for reliable index tuning. We propose Beauty, the first uncertainty-aware framework that enhances learning-based models with uncertainty quantification and uses what-if tools as a complementary mechanism to improve reliability and reduce management complexity. Specifically, we introduce a novel method that combines AutoEncoder and Monte Carlo Dropout to jointly quantify uncertainty, tailored to the characteristics of benefit estimation tasks. In experiments involving sixteen models, our approach outperformed existing uncertainty quantification methods in the majority of cases. We also conducted index tuning tests on six datasets. By applying the Beauty framework, we eliminated worst-case scenarios and more than tripled the occurrence of best-case scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。