arXiv:2506.13172cs.CLcs.AI2025-06

用提示词让AI识别论文摘要中的无依据结论和模糊代词,提升学术文本可信度。

AI-Facilitated Analysis of Abstracts and Conclusions: Flagging Unsubstantiated Claims and Ambiguous Pronouns

  • 设计结构化提示词,引导AI进行层次化语义与语言分析。
  • 模型对名词短语中无依据成分识别率达95%,但修饰语识别差异显著。
  • 上下文完整时表现良好,摘要单独分析时性能波动大,需针对性测试。

我们提出并评估了一套概念验证(PoC)级的结构化提示流程,旨在引导大型语言模型(LLMs)在学术稿件的高层语义与语言分析中展现类人层级推理能力。该流程聚焦于摘要与结论部分的两项非平凡分析任务:识别无依据陈述(信息完整性)与标记语义模糊的歧义代词引用(语言清晰性)。我们在两个前沿模型(Gemini Pro 2.5 Pro 和 ChatGPT Plus o3)上进行了多轮系统性评估,考察不同上下文条件下的表现。在信息完整性任务中,两模型均能成功识别名词短语中的无依据主干(95% 成功率),但 ChatGPT 在识别无依据形容词修饰语时完全失败(0% 成功率),而 Gemini 正确识别率达 95%,提示句法角色可能影响表现。在语言分析任务中,两者在完整论文上下文中表现良好(80–90% 成功率)。然而,在仅提供摘要的条件下,Gemini 性能显著下降,而 ChatGPT 达到 100% 成功率。结果表明,虽然结构化提示是复杂文本分析的有效方法,但其效果高度依赖模型、任务类型与上下文的交互,强调需开展模型特定的严谨测试。

原文摘要 · Abstract (English)

We present and evaluate a suite of proof-of-concept (PoC), structured workflow prompts designed to elicit human-like hierarchical reasoning while guiding Large Language Models (LLMs) in the high-level semantic and linguistic analysis of scholarly manuscripts. The prompts target two non-trivial analytical tasks within academic summaries (abstracts and conclusions): identifying unsubstantiated claims (informational integrity) and flagging semantically confusing ambiguous pronoun references (linguistic clarity). We conducted a systematic, multi-run evaluation on two frontier models (Gemini Pro 2.5 Pro and ChatGPT Plus o3) under varied context conditions. Our results for the informational integrity task reveal a significant divergence in model performance: while both models successfully identified an unsubstantiated head of a noun phrase (95% success), ChatGPT consistently failed (0% success) to identify an unsubstantiated adjectival modifier that Gemini correctly flagged (95% success), raising a question regarding the potential influence of the target's syntactic role. For the linguistic analysis task, both models performed well (80-90% success) with full manuscript context. Surprisingly, in a summary-only setting, Gemini's performance was substantially degraded, while ChatGPT achieved a perfect (100%) success rate. Our findings suggest that while structured prompting is a viable methodology for complex textual analysis, prompt performance may be highly dependent on the interplay between the model, task type, and context, highlighting the need for rigorous, model-specific testing.

AI辅助文本分析学术写作提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。