arXiv:2505.12257cs.CYcs.AI2025-05被引 1

用提示工程让大模型更精准发现化学公式的图文错误。

LLM Context Conditioning and PWP Prompting for Multimodal Validation of Chemical Formulas

  • 通过结构化提示控制大模型分析思路,避免其自动纠错掩盖错误。
  • 该方法使Gemini 2.5 Pro发现人眼遗漏的图像公式错误,而ChatGPT失败。
  • 仅用标准聊天界面即可实现,无需改模型或调API,适合科研人员使用。

识别复杂科学文档中需多模态解读(如图像中的化学公式)的细微技术错误,对大语言模型(LLMs)构成重大挑战,因其固有的纠错倾向可能掩盖真实错误。本探索性概念验证(PoC)研究考察了基于持久工作流提示(PWP)原则的结构化上下文条件化方法,作为在推理阶段调节LLM行为的策略。该方法旨在提升通用型大模型(如Gemini 2.5 Pro和ChatGPT Plus o3)在精确验证任务中的可靠性,仅依赖标准聊天界面,无需API访问或模型修改。研究聚焦于一篇已知存在文本与图像错误的复杂论文,评估多种提示策略:基础提示不可靠,而采用PWP结构严格引导分析思维的方法显著提升了两种模型对文本错误的识别能力。尤为关键的是,该方法成功引导Gemini 2.5 Pro反复识别出此前人工审查遗漏的图像公式错误,而ChatGPT Plus o3在测试中未能完成此任务。初步结果揭示了阻碍细致验证的特定模型行为模式,并表明受PWP启发的上下文条件化是一种极具潜力且高可及性的技术,可用于构建更稳健的基于大模型的分析流程,尤其适用于科学与技术文档中需要精密纠错的任务。未来需在更大范围验证其普适性。

原文摘要 · Abstract (English)

Identifying subtle technical errors within complex scientific and technical documents, especially those requiring multimodal interpretation (e.g., formulas in images), presents a significant hurdle for Large Language Models (LLMs) whose inherent error-correction tendencies can mask inaccuracies. This exploratory proof-of-concept (PoC) study investigates structured LLM context conditioning, informed by Persistent Workflow Prompting (PWP) principles, as a methodological strategy to modulate this LLM behavior at inference time. The approach is designed to enhance the reliability of readily available, general-purpose LLMs (specifically Gemini 2.5 Pro and ChatGPT Plus o3) for precise validation tasks, crucially relying only on their standard chat interfaces without API access or model modifications. To explore this methodology, we focused on validating chemical formulas within a single, complex test paper with known textual and image-based errors. Several prompting strategies were evaluated: while basic prompts proved unreliable, an approach adapting PWP structures to rigorously condition the LLM's analytical mindset appeared to improve textual error identification with both models. Notably, this method also guided Gemini 2.5 Pro to repeatedly identify a subtle image-based formula error previously overlooked during manual review, a task where ChatGPT Plus o3 failed in our tests. These preliminary findings highlight specific LLM operational modes that impede detail-oriented validation and suggest that PWP-informed context conditioning offers a promising and highly accessible technique for developing more robust LLM-driven analytical workflows, particularly for tasks requiring meticulous error detection in scientific and technical documents. Extensive validation beyond this limited PoC is necessary to ascertain broader applicability.

大模型化学信息提示工程错误检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。