arXiv:2603.13173cs.AIcs.CL2026-03中稿 · publication in 20t…被引 3

测试大模型在语义不变输入下的推理稳定性,发现小模型反而更可靠。

Semantic Invariance in Agentic AI

  • 设计八种语义保持变换,系统评估大模型推理鲁棒性
  • 小模型Qwen3-30B-A3B稳定率达79.6%,大于模型更易波动
  • 适用于评估高风险场景中AI代理的可靠性,如科学决策

大语言模型日益作为自主推理代理用于决策支持、科学问题求解和多智能体协同系统。然而,在关键应用中部署时,需确保其推理在语义等价输入变化下保持稳定,这一特性称为语义不变性。现有标准评测仅针对固定问题形式,无法捕捉此关键可靠性维度。为此,本文提出一种元测试框架,系统评估大模型推理的鲁棒性,对七种基础模型(涵盖四种架构家族:Hermes 70B/405B、Qwen3 30B-A3B/235B-A22B、DeepSeek-R1、gpt-oss 20B/120B)应用八种语义保持变换(恒等、改写、事实重排、扩展、压缩、学术语境、商业语境、对比表述),覆盖19个跨八大学科的多步推理问题。结果表明,模型规模无法预测鲁棒性:较小的Qwen3-30B-A3B达到最高稳定性(79.6%不变响应,语义相似度0.91),而更大模型表现更脆弱。

原文摘要 · Abstract (English)

Large Language Models (LLMs) increasingly serve as autonomous reasoning agents in decision support, scientific problem-solving, and multi-agent coordination systems. However, deploying LLM agents in consequential applications requires assurance that their reasoning remains stable under semantically equivalent input variations, a property we term semantic invariance. Standard benchmark evaluations, which assess accuracy on fixed, canonical problem formulations, fail to capture this critical reliability dimension. To address this shortcoming, in this paper we present a metamorphic testing framework for systematically assessing the robustness of LLM reasoning agents, applying eight semantic-preserving transformations (identity, paraphrase, fact reordering, expansion, contraction, academic context, business context, and contrastive formulation) across seven foundation models spanning four distinct architectural families: Hermes (70B, 405B), Qwen3 (30B-A3B, 235B-A22B), DeepSeek-R1, and gpt-oss (20B, 120B). Our evaluation encompasses 19 multi-step reasoning problems across eight scientific domains. The results reveal that model scale does not predict robustness: the smaller Qwen3-30B-A3B achieves the highest stability (79.6% invariant responses, semantic similarity 0.91), while larger models exhibit greater fragility.

大模型推理语义不变性鲁棒性测试AI代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。