arXiv:2510.00300cs.AI2025-10

ICL引导提升效率却削弱复杂推理能力,导致系统性脆弱。

ICL Optimized Fragility

  • 用六种ICL配置测试大模型跨领域推理表现
  • 常识题准确率91%-99%,逻辑谜题降至10%-43%
  • 数学奥赛题不受影响,适合对效率敏感场景

ICL引导已知可提升特定任务性能,但其对跨领域认知能力的影响尚不明确。本研究使用六种GPT-OSS:20b模型变体(一种基线模型及五种ICL配置:简单、思维链、随机、附加文本、符号语言)在840项测试中评估其表现,涵盖常识问题、逻辑谜题与数学奥赛题。统计分析(ANOVA)显示各ICL变体间存在显著行为差异(p<0.001),揭示“优化脆弱性”现象。模型在常识任务中准确率达91%-99%,但在复杂推理任务中表现下降,谜题准确率降至10%-43%(基线为43%)。数学奥赛题无显著差异(p=0.2173),表明复杂数学推理不受影响。结果表明ICL引导带来效率与推理灵活性的系统性权衡,对大模型部署与AI安全具有重要启示。

原文摘要 · Abstract (English)

ICL guides are known to improve task-specific performance, but their impact on cross-domain cognitive abilities remains unexplored. This study examines how ICL guides affect reasoning across different knowledge domains using six variants of the GPT-OSS:20b model: one baseline model and five ICL configurations (simple, chain-of-thought, random, appended text, and symbolic language). The models were subjected to 840 tests spanning general knowledge questions, logic riddles, and a mathematical olympiad problem. Statistical analysis (ANOVA) revealed significant behavioral modifications (p less than 0.001) across ICL variants, demonstrating a phenomenon termed "optimized fragility." ICL models achieved 91%-99% accuracy on general knowledge tasks while showing degraded performance on complex reasoning problems, with accuracy dropping to 10-43% on riddles compared to 43% for the baseline model. Notably, no significant differences emerged on the olympiad problem (p=0.2173), suggesting that complex mathematical reasoning remains unaffected by ICL optimization. These findings indicate that ICL guides create systematic trade-offs between efficiency and reasoning flexibility, with important implications for LLM deployment and AI safety.

ICL推理脆弱性大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。