测试聊天机器人在金融文献中的引用准确性,发现其幻觉率高达76.7%。
Evaluating the Accuracy of Chatbots in Financial Literature
- 设计非二元评估方法与时效性指标,分析幻觉随主题新旧变化
- ChatGPT-4o幻觉率为20.0%,Gemini Advanced达76.7%
- 适用于需要验证引用可靠性的金融研究者
我们评估了两个聊天机器人(ChatGPT 4o 和 o1-preview 版本)以及 Gemini Advanced 在提供金融文献参考方面的可靠性,并采用新颖的方法学。除了文献中常见的二元判断方式外,我们还开发了非二元评估方法和时效性指标,以分析幻觉率如何随主题的新旧程度变化。通过对150个引用的分析,ChatGPT-4o 的幻觉率为20.0%(95%置信区间:13.6%-26.4%),o1-preview 为21.3%(95%置信区间:14.8%-27.9%)。相比之下,Gemini Advanced 的幻觉率更高,达到76.7%(95%置信区间:69.9%-83.4%)。虽然幻觉率在较新主题中有所上升,但该趋势对 Gemini Advanced 并无统计显著性。研究强调了在快速演变领域中验证聊天机器人所提供参考的重要性。
原文摘要 · Abstract (English)
We evaluate the reliability of two chatbots, ChatGPT (4o and o1-preview versions), and Gemini Advanced, in providing references on financial literature and employing novel methodologies. Alongside the conventional binary approach commonly used in the literature, we developed a nonbinary approach and a recency measure to assess how hallucination rates vary with how recent a topic is. After analyzing 150 citations, ChatGPT-4o had a hallucination rate of 20.0% (95% CI, 13.6%-26.4%), while the o1-preview had a hallucination rate of 21.3% (95% CI, 14.8%-27.9%). In contrast, Gemini Advanced exhibited higher hallucination rates: 76.7% (95% CI, 69.9%-83.4%). While hallucination rates increased for more recent topics, this trend was not statistically significant for Gemini Advanced. These findings emphasize the importance of verifying chatbot-provided references, particularly in rapidly evolving fields.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。