arXiv:2512.22508cs.LGcs.AI2025-12中稿 · as a Short Paper a…

用元数据和幻觉信号预测LLM在牙科题中的答案正确性。

Predicting LLM Correctness in Prosthodontics Using Metadata and Hallucination Signals

  • 基于元数据和幻觉信号构建正确性预测模型。
  • 准确率最高提升7.14%,精度达83.12%。
  • 适合关注LLM可靠性评估的研究者。

大语言模型(LLMs)在医疗、医学教育等高风险领域应用日益广泛,但生成事实错误(即幻觉)信息的风险令人担忧。尽管已有大量研究致力于检测和缓解幻觉,但预测模型回答是否正确仍是关键却未被充分探索的问题。本研究考察了通用模型GPT-4o与推理导向模型OSS-120B在多项选择牙科考试中的表现,通过三种不同提示策略分析元数据和幻觉信号,为每对(模型,提示)构建正确性预测器。结果表明,该元数据方法可使准确率最高提升7.14%,在基准假设所有答案正确的情况下达到83.12%的精度。研究还发现,实际幻觉是错误的强指标,但元数据本身无法可靠预测幻觉。此外,尽管提示策略不影响整体准确率,却显著改变模型内部行为及元数据的预测效用。这些结果为开发大模型可靠性信号提供了前景,但也表明当前方法尚不足以支持高风险场景部署。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly adopted in high-stakes domains such as healthcare and medical education, where the risk of generating factually incorrect (i.e., hallucinated) information is a major concern. While significant efforts have been made to detect and mitigate such hallucinations, predicting whether an LLM's response is correct remains a critical yet underexplored problem. This study investigates the feasibility of predicting correctness by analyzing a general-purpose model (GPT-4o) and a reasoning-centric model (OSS-120B) on a multiple-choice prosthodontics exam. We utilize metadata and hallucination signals across three distinct prompting strategies to build a correctness predictor for each (model, prompting) pair. Our findings demonstrate that this metadata-based approach can improve accuracy by up to +7.14% and achieve a precision of 83.12% over a baseline that assumes all answers are correct. We further show that while actual hallucination is a strong indicator of incorrectness, metadata signals alone are not reliable predictors of hallucination. Finally, we reveal that prompting strategies, despite not affecting overall accuracy, significantly alter the models' internal behaviors and the predictive utility of their metadata. These results present a promising direction for developing reliability signals in LLMs but also highlight that the methods explored in this paper are not yet robust enough for critical, high-stakes deployment.

大模型幻觉检测可靠性评估牙科AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。