揭示大模型如何用不同线索编码反问句的修辞意图
Rhetorical Questions in LLM Representations: A Linear Probing Study
- 用线性探测分析社交媒体数据中的反问句表示
- 反问句在最后词表示中稳定可辨,跨数据集检测AUROC达0.7-0.8
- 不同数据训练的探针捕捉不同修辞现象,无统一表征
反问句用于说服或表达立场而非获取信息。大语言模型如何内部表征反问句尚不明确。我们在两个具有不同话语语境的社会媒体数据集上,使用线性探测分析大模型对反问句的表征,发现反问信号早期出现,并由最后词表示最稳定地捕获。在同一数据集中,反问句与信息寻求句呈线性可分,且在跨数据集迁移下仍可检测,AUROC约为0.7–0.8。然而,我们证明迁移能力并不意味着共享表征:在相同目标语料上应用时,不同数据集训练的探针产生不同的实例排名,前几名重叠率常低于0.2。定性分析显示,这些差异对应于不同修辞现象:部分探针捕捉长程论证中的话语级立场,另一些则聚焦局部语法驱动的疑问行为。结果表明,大模型中反问句的表征由多个强调不同线索的线性方向构成,而非单一共享方向。
原文摘要 · Abstract (English)
Rhetorical questions are asked not to seek information but to persuade or signal stance. How large language models internally represent them remains unclear. We analyze rhetorical questions in LLM representations using linear probes on two social-media datasets with different discourse contexts, and find that rhetorical signals emerge early and are most stably captured by last-token representations. Rhetorical questions are linearly separable from information-seeking questions within datasets, and remain detectable under cross-dataset transfer, reaching AUROC around 0.7-0.8. However, we demonstrate that transferability does not simply imply a shared representation. Probes trained on different datasets produce different rankings when applied to the same target corpus, with overlap among the top-ranked instances often below 0.2. Qualitative analysis shows that these divergences correspond to distinct rhetorical phenomena: some probes capture discourse-level rhetorical stance embedded in extended argumentation, while others emphasize localized, syntax-driven interrogative acts. Together, these findings suggest that rhetorical questions in LLM representations are encoded by multiple linear directions emphasizing different cues, rather than a single shared direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。