arXiv:2602.22359cs.CLcs.AI2026-02

用GPT-5分析引文语境,发现提示词设计直接影响解读结果。

Scaling In, Not Up? Testing Thick Citation Context Analysis with GPT-5 and Fragile Prompts

  • 通过两阶段流程解析引文,结合文本与上下文重构意义
  • 生成450条假设,识别出21种常见解读模式
  • 提示词设计能系统性影响模型关注点和表达词汇

本研究测试大语言模型(LLMs)在厚文本语境下进行引文分析(CCA)的可行性,不依赖类型标签的扩展,而是聚焦单个复杂案例的深度阅读。采用平衡的2×3实验设计,改变提示框架与支撑结构,以Chubin和Moitra(1975)第6条脚注及Gilbert(1977)重构为分析样本,构建两阶段GPT-5流程:第一阶段仅基于引文文本进行表面分类与预期判断,第二阶段结合引用文献与被引文献全文进行跨文档解释重构。共完成90次重构,生成450条不同假设。通过细致阅读与归纳编码,识别出21种反复出现的解释策略;线性概率模型显示提示选择显著影响这些策略的频率与词汇范围。GPT-5在表面阶段高度稳定,始终将该引文归类为“补充性”。在重构阶段,模型生成了结构化的合理解读空间,但提示的支架与示例会转移注意力并改变用词,有时导向牵强的解读。相较Gilbert,GPT-5虽捕捉到相同的文本关键点,但更倾向于将其解释为谱系关系与定位,而非警告。研究揭示了将LLM作为可检查、可争议的引导式合作者在解释性引文分析中的潜力与风险,并表明提示设计会系统性地塑造模型所强调的可能解读与表达方式。

原文摘要 · Abstract (English)

This paper tests whether large language models (LLMs) can support interpretative citation context analysis (CCA) by scaling in thick, text-grounded readings of a single hard case rather than scaling up typological labels. It foregrounds prompt-sensitivity analysis as a methodological issue by varying prompt scaffolding and framing in a balanced 2x3 design. Using footnote 6 in Chubin and Moitra (1975) and Gilbert's (1977) reconstruction as a probe, I implement a two-stage GPT-5 pipeline: a citation-text-only surface classification and expectation pass, followed by cross-document interpretative reconstruction using the citing and cited full texts. Across 90 reconstructions, the model produces 450 distinct hypotheses. Close reading and inductive coding identify 21 recurring interpretative moves, and linear probability models estimate how prompt choices shift their frequencies and lexical repertoire. GPT-5's surface pass is highly stable, consistently classifying the citation as "supplementary". In reconstruction, the model generates a structured space of plausible alternatives, but scaffolding and examples redistribute attention and vocabulary, sometimes toward strained readings. Relative to Gilbert, GPT-5 detects the same textual hinges yet more often resolves them as lineage and positioning than as admonishment. The study outlines opportunities and risks of using LLMs as guided co-analysts for inspectable, contestable interpretative CCA, and it shows that prompt scaffolding and framing systematically tilt which plausible readings and vocabularies the model foregrounds.

引文分析提示工程大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。