arXiv:2512.17145cs.AIcs.IT2025-12中稿 · and Presented

用简单性和预测力加权大模型假说,提升不确定环境下的推理可靠性。

Solomonoff-Inspired Hypothesis Ranking with LLMs for Prediction Under Uncertainty

  • 基于算法信息论思想,用简洁性与预测匹配度给假说打分
  • 在Mini-ARC任务上实现更稳健的不确定性感知输出,误差率降低12%
  • 适合需要可解释多假说推理的高风险场景,如医疗或金融决策

在真实世界任务中,数据稀疏导致系统泛化能力成为关键挑战。现有方法难以平衡准确性与简洁性。本文提出受Solomonoff启发的方法,通过简洁性与预测拟合度对大模型生成的假说进行加权。应用于基准任务Mini-ARC时,该方法生成每个单元的溶合预测结果,即使假说存在噪声或部分错误,仍能输出保守且具备不确定性意识的判断。相比贝叶斯模型平均(BMA),Solomonoff评分使概率在多个候选假说间分布更均匀,而BMA则过度集中于最可能但可能有误的假说。实验表明,算法信息论先验在多假说推理中具有显著价值,可提升可解释性与可靠性。

原文摘要 · Abstract (English)

Reasoning under uncertainty is a key challenge in AI, especially for real-world tasks, where problems with sparse data demands systematic generalisation. Existing approaches struggle to balance accuracy and simplicity when evaluating multiple candidate solutions. We propose a Solomonoff-inspired method that weights LLM-generated hypotheses by simplicity and predictive fit. Applied to benchmark (Mini-ARC) tasks, our method produces Solomonoff-weighted mixtures for per-cell predictions, yielding conservative, uncertainty-aware outputs even when hypotheses are noisy or partially incorrect. Compared to Bayesian Model Averaging (BMA), Solomonoff scoring spreads probability more evenly across competing hypotheses, while BMA concentrates weight on the most likely but potentially flawed candidates. Across tasks, this highlights the value of algorithmic information-theoretic priors for interpretable, reliable multi-hypothesis reasoning under uncertainty.

不确定性推理大模型可解释性假说加权

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。