arXiv:2509.20750cs.CLcs.AI2025-09EMNLP被引 3

用自信度引导重构问题,零样本问答更准更稳。

Confidence-guided Refinement Reasoning for Zero-shot Question Answering

  • 通过构建并优化子问题与答案,动态提升目标答案的置信度。
  • 在多个模型和基准上实现稳定性能提升,无需重新训练。
  • 揭示子问题数量与质量对推理可靠性的影响,指导实践。

我们提出了一种无需训练的新型框架C2R,适用于文本、图像和视频等多模态问答任务。C2R通过策略性构建和精炼子问题及其答案(子QA),提升目标答案的置信度评分。首先,从子QA中筛选出多样化推理路径;随后,比较不同候选答案的置信度,选择最可靠的最终答案。由于C2R仅依赖模型自身输出的置信度,可无缝集成到多种现有问答模型中,并在多个模型和基准上持续提升性能。此外,我们提供了关于子QA利用如何影响模型行为的重要但被忽视的见解,具体分析了子QA的数量与质量对实现鲁棒可靠推理的影响。

原文摘要 · Abstract (English)

We propose Confidence-guided Refinement Reasoning (C2R), a novel training-free framework applicable to question-answering (QA) tasks across text, image, and video domains. C2R strategically constructs and refines sub-questions and their answers (sub-QAs), deriving a better confidence score for the target answer. C2R first curates a subset of sub-QAs to explore diverse reasoning paths, then compares the confidence scores of the resulting answer candidates to select the most reliable final answer. Since C2R relies solely on confidence scores derived from the model itself, it can be seamlessly integrated with various existing QA models, demonstrating consistent performance improvements across diverse models and benchmarks. Furthermore, we provide essential yet underexplored insights into how leveraging sub-QAs affects model behavior, specifically analyzing the impact of both the quantity and quality of sub-QAs on achieving robust and reliable reasoning.

零样本问答置信度引导推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。