将置信预测扩展到图文多模态回归,实现无分布假设的不确定性量化。
Conformal Prediction for Multimodal Regression
- 利用神经网络融合点的内部特征构建预测区间
- 在多模态数据上实现可靠且无需分布假设的不确定性估计
- 适合需要可信预测区间的医疗、自动驾驶等多模态场景
本文提出多模态置信预测。传统置信预测仅适用于纯数值输入,现通过新方法拓展至处理图像和非结构化文本的复杂神经网络架构。研究发现,从多模态信息融合的收敛点提取的内部特征,可有效用于构建预测区间(PIs)。该能力为富含多模态数据的领域开辟了新路径,使更多问题能受益于无分布假设的不确定性量化。
原文摘要 · Abstract (English)
This paper introduces multimodal conformal regression. Traditionally confined to scenarios with solely numerical input features, conformal prediction is now extended to multimodal contexts through our methodology, which harnesses internal features from complex neural network architectures processing images and unstructured text. Our findings highlight the potential for internal neural network features, extracted from convergence points where multimodal information is combined, to be used by conformal prediction to construct prediction intervals (PIs). This capability paves new paths for deploying conformal prediction in domains abundant with multimodal data, enabling a broader range of problems to benefit from guaranteed distribution-free uncertainty quantification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。