测试大模型对语序敏感度,发现恢复结构仍很困难。
How Order-Sensitive Are LLMs? OrderProbe for Deterministic Structural Reconstruction
- 用中日韩四字固定表达构建确定性评测基准
- 顶尖模型零样本恢复率低于35%,结构重建仍难
- 揭示语义理解与结构恢复能力不一致,适合研究模型认知机制
大型语言模型在语义理解方面表现优异,但其从乱序输入中重构内部结构的能力仍缺乏系统探索。句子级恢复难以自动评估,因为乱序句子常存在多个有效重排方式。本文提出OrderProbe,一种基于中文、日文和韩文固定四字表达的确定性结构重建评测基准,这些表达具有唯一标准顺序,支持精确匹配评分。同时构建诊断框架,评估模型在恢复准确率之外的表现,包括语义准确性、逻辑有效性、结构一致性、鲁棒性和信息密度。在12个主流LLM上的实验表明,即使前沿模型在结构重建上仍面临挑战:零样本恢复率经常低于35%。此外,我们观察到以意义为导向的生成与精确结构重建之间存在持续差距,表明结构鲁棒性并非语义能力的自然产物。
原文摘要 · Abstract (English)
Large language models (LLMs) excel at semantic understanding, yet their ability to reconstruct internal structure from scrambled inputs remains underexplored. Sentence-level restoration is difficult to evaluate automatically because scrambled sentences often admit multiple valid reorderings. We introduce OrderProbe, a deterministic benchmark for structural reconstruction using fixed four-character expressions in Chinese, Japanese, and Korean, which have a unique canonical order and thus support exact-match scoring. We further propose a diagnostic framework that evaluates models beyond recovery accuracy, including Semantic Accuracy, Logical Validity, Structural Consistency, Robustness, and Information Density. Experiments on twelve widely used LLMs show that structural reconstruction remains difficult even for frontier systems: zero-shot recovery frequently falls below 35%. We also observe a consistent gap between meaning-oriented generation and exact structural reconstruction, suggesting that structural robustness is not an automatic byproduct of semantic competence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。