arXiv:2412.13645cs.AIcs.CL2024-12EMNLP

模型内在先验比示范数据更能决定推理质量,即使没有示范也能保持良好表现。

On the Role of Model Prior in Real-World Inductive Reasoning

  • 通过对比三种推理策略,发现模型先验主导假设生成
  • 移除示范后假设质量仅轻微下降,下游任务表现基本不变
  • 先验难以被示范或标签翻转覆盖,适合真实场景应用

大型语言模型(LLMs)在归纳推理方面表现出色,能根据上下文示范生成可泛化的假设。然而在真实应用中,假设生成不仅依赖示范,更受任务特定模型先验的显著影响。尽管其作用关键,但模型先验与示范各自的贡献尚未被充分研究。本研究通过在五个真实任务上对三类模型进行系统评估,发现假设生成主要由模型内在先验驱动;移除示范仅导致假设质量轻微下降,下游使用效果几乎不受影响。进一步分析表明,该结果在不同标签格式和配置下均一致,且先验难以被颠覆,即便在标签反转情况下仍保持稳定。这些发现深化了对大模型假设生成机制的理解,并提示应更有效地利用模型先验以提升实际推理性能。

原文摘要 · Abstract (English)

Large Language Models (LLMs) show impressive inductive reasoning capabilities, enabling them to generate hypotheses that could generalize effectively to new instances when guided by in-context demonstrations. However, in real-world applications, LLMs' hypothesis generation is not solely determined by these demonstrations but is significantly shaped by task-specific model priors. Despite their critical influence, the distinct contributions of model priors versus demonstrations to hypothesis generation have been underexplored. This study bridges this gap by systematically evaluating three inductive reasoning strategies across five real-world tasks with three LLMs. Our empirical findings reveal that, hypothesis generation is primarily driven by the model's inherent priors; removing demonstrations results in minimal loss of hypothesis quality and downstream usage. Further analysis shows the result is consistent across various label formats with different label configurations, and prior is hard to override, even under flipped labeling. These insights advance our understanding of the dynamics of hypothesis generation in LLMs and highlight the potential for better utilizing model priors in real-world inductive reasoning tasks.

大模型推理模型先验归纳推理真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。