提出可生成语义多样的数学题变体的新框架,减少模型记忆测试集的干扰。
GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations

- 通过修改实体、属性和关系生成语义上差异大的题目变体
- 14个顶尖大模型在新基准上平均下降28%,证明其挑战性
- 适用于数学、逻辑等多类任务,释放全人工验证数据集
GSM8K等基准常被用作数学推理能力的衡量标准,但排行榜提升可能因模型记忆固定测试集而夸大真实能力。现有鲁棒性测试多为表面扰动(改写、重命名、数字互换、干扰项),基本保留原始事实,静态发布版本本身也可能成为记忆目标。我们提出GSM-SEM,一种可重复使用且随机生成的语义多样化基准变体框架,其生成的变体在语义上显著区别于原有题目。该方法通过修改问题中的实体、属性或关系,频繁改变底层事实,迫使模型在新条件下重新计算答案,同时保持原题的计算过程与近似难度。每次运行均可自动生成新变体,无需重新标注,降低对静态公开基准的依赖,从而减轻记忆偏差。我们在GSM8K及两个已有变体集(GSM-Symbolic和GSM-Plus)上应用GSM-SEM,生成GSM8K-SEM、GSM-Symbolic-SEM和GSM-Plus-SEM。评估14个SOTA大模型发现,性能普遍下降,尤其当语义扰动与符号/附加变化结合时下降更明显(最大严格度下平均下降28%)。我们公开发布这三个经人工验证的变体数据集。最后,为展示其泛化能力,我们将GSM-SEM应用于BigBenchHard、LogicBench和NLR-BIRD等其他基准。
原文摘要 · Abstract (English)
Benchmarks like GSM8K are popular measures of mathematical reasoning, but leaderboard gains can overstate true capability due to memorization of fixed test sets. Most robustness variants apply surface-level perturbations (paraphrases, renamings, number swaps, distractors) that largely preserve the underlying facts, and static releases can themselves become memorization targets over time. We introduce GSM-SEM, a reusable and stochastic framework for generating semantically diverse benchmark variants with substantially higher semantic variance than prior approaches. GSM-SEM perturbs problem statements by modifying entities, attributes, and/or relationships, frequently altering underlying facts and requiring models to recompute solutions under new conditions, while constraining generation to preserve the original calculations/answer and approximate problem difficulty. GSM-SEM generates fresh variants on each run without requiring re-annotation, reducing reliance on static public benchmarks for evaluation and thereby lowering the bias of memorization. We apply GSM-SEM on GSM8K and two existing variation suites (GSM-Symbolic and GSM-Plus), producing GSM8K-SEM, GSM-Symbolic-SEM, and GSM-Plus-SEM. Evaluating 14 SOTA LLMs, we observe consistent performance drops with larger decline when semantic perturbations are coupled with symbolic/plus variations (average drop rate 28% in maximum strictness configuration of GSM-SEM). We publicly release the three SEM variants as fully human-validated datasets. Finally, to demonstrate applicability beyond GSM-style math problems, we apply GSM-SEM to additional benchmarks including BigBenchHard, LogicBench, and NLR-BIRD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。