GPT-5能自动纠正GPT-4V读图错误,无需额外提示。
GPT-5 Model Corrected GPT-4V's Chart Reading Errors, Not Prompting
- 用GPT-5直接修正GPT-4V在复杂图像上的读图错误。
- 在107个可视化问题中,GPT-5准确率显著高于GPT-4V。
- 模型架构比提示词设计对准确率影响更大,适合做图表分析研究者参考。
我们通过定量评估,研究零样本大语言模型(LLMs)和提示词使用对图表阅读任务的影响。要求模型回答107个可视化问题,比较代理型GPT-5与多模态GPT-4V在困难图像实例上的推理准确性,这些情况下GPT-4V无法给出正确答案。结果显示,模型架构主导推理准确率:GPT-5大幅提升了准确率,而提示词变体仅带来微小改善。本研究预注册信息见https://osf.io/u78td/?view_only=6b075584311f48e991c39335c840ded3;Google Drive材料链接为https://drive.google.com/file/d/1ll8WWZDf7cCNcfNWrLViWt8GwDNSvVrp/view。
原文摘要 · Abstract (English)
We present a quantitative evaluation to understand the effect of zero-shot large-language model (LLMs) and prompting uses on chart reading tasks. We asked LLMs to answer 107 visualization questions to compare inference accuracies between the agentic GPT-5 and multimodal GPT-4V, for difficult image instances, where GPT-4V failed to produce correct answers. Our results show that model architecture dominates the inference accuracy: GPT5 largely improved accuracy, while prompt variants yielded only small effects. Pre-registration of this work is available here: https://osf.io/u78td/?view_only=6b075584311f48e991c39335c840ded3; the Google Drive materials are here:https://drive.google.com/file/d/1ll8WWZDf7cCNcfNWrLViWt8GwDNSvVrp/view.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。