提出复杂度指标SICI,揭示大模型立场判断的失效规律。
SICI: A Semantic-Pragmatic Complexity Index Reveals Regime Shifts in LLM Stance Detection

- 用七维指标SICI量化文本与立场间的语义语用负担。
- 高复杂度下模型错误从误判转向放弃判断,呈现相变特征。
- 适用于研究大模型在复杂推理中的瓶颈与改进方向。
基于提示的大语言模型广泛用于立场检测,但更难的例子未必能通过更清晰指令、推理提示、检索或辩论修复。本文提出SICI(立场推断复杂度指数),一个七维诊断指标,用于衡量目标-文本对带来的语义语用负担。在SemEval-2016和VAST数据集上,SICI对模型准确率的预测优于表面代理指标,且跨评分者一致性较高(α=0.771)。更重要的是,随着SICI升高,模型错误模式发生范式转变:低复杂度样本易导致过度归因,尤其在“反对”类判断中;中等复杂度样本形成不稳定的边界;高复杂度样本则迅速集中于“无立场”判断。这种类似相变的结构在GPT-3.5、GPT-4o-mini、DeepSeek-V3和GPT-4o中均持续存在,尽管更强模型会将边界向右移动。15种干预方法实验进一步表明,提示、检索和辩论通常仅沿归因-回避轴调整模型行为,而非突破高复杂度瓶颈。
原文摘要 · Abstract (English)
Prompt-based LLMs are increasingly used for stance detection, but harder examples are not always repaired by clearer instructions, reasoning prompts, retrieval, or debate. We introduce SICI (Stance Inference Complexity Index), a seven-dimensional diagnostic measure of the semantic-pragmatic burden imposed by a target--text pair. Across SemEval-2016 and VAST, SICI predicts LLM accuracy better than surface proxies and shows substantial cross-scorer reliability ($α=0.771$). More importantly, LLM errors change regime as SICI increases: low-complexity examples invite over-attribution, especially Against predictions; intermediate examples form an unstable boundary; and high-complexity examples rapidly concentrate on None. This phase-transition-like structure persists across GPT-3.5, GPT-4o-mini, DeepSeek-V3, and GPT-4o, although stronger models move the boundaries. A 15-method intervention study further shows that prompting, retrieval, and debate often shift models along the attribution--abstention axis rather than removing the high-complexity bottleneck.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。