arXiv:2604.17794cs.CLcs.AI2026-04

小模型在越南语数学推理中表现不佳,通过微调和简单测试时扩展可显著提升

Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling

  • 用微调激活小模型潜在知识,解决越南语推理中的表达缺陷
  • 微调后解释质量提升77%,计算与教学连贯性明显改善
  • 简单链式思考加自一致比复杂推理框架更适合边缘设备

小型语言模型(SLMs)在资源受限设备上部署复杂推理能力面临挑战,尤其在非英语语言如越南语中。本文研究基于Qwen3-1.7B架构的测试时扩展策略,聚焦越南语初等数学任务。构建了高保真推理数据集Vi-S1K和双资源评估基准Vi-Elementary-Bench,采用大模型评判协议发现,基础模型具备强潜在知识(准确率4.05/5.00),但存在严重“格式缺口”。监督微调(SFT)作为关键“推理解锁器”,使解释质量提升77%,弥合计算与教学连贯性差距。进一步分析显示,结构化框架如ReAct对1.7B参数模型造成“认知负担”,性能低于纯链式思考(CoT)结合自一致性策略。结果确立了小模型部署层级:微调配合简化测试时扩展优于复杂智能体流程。

原文摘要 · Abstract (English)

The democratization of ubiquitous AI hinges on deploying sophisticated reasoning capabilities on resource-constrained devices. However, Small Language Models (SLMs) often face a "reasoning gap", particularly in non-English languages like Vietnamese, where they struggle to maintain coherent chains of thought. This paper investigates Test-Time Scaling strategies for the Qwen3-1.7B architecture within the context of Vietnamese Elementary Mathematics. We introduce Vi-S1K, a high-fidelity reasoning dataset localized via a Gemini 2.5 Flash-Lite powered pipeline, and Vi-Elementary-Bench, a dual-resource benchmark for rigorous evaluation. Using an LLM-as-a-Judge protocol, we reveal that the base model possesses robust latent knowledge (Accuracy: 4.05/5.00) but suffers from a severe "formatting gap" in communication. Supervised Fine-Tuning (SFT) acts as a critical "reasoning unlocker", yielding a 77% improvement in Explanation Quality and bridging the gap between raw calculation and pedagogical coherence. Furthermore, our analysis of prompting strategies uncovers a significant trade-off: structured frameworks like ReAct impose a "cognitive tax" on the 1.7B parameter capacity, degrading performance relative to pure Chain-of-Thought (CoT) combined with Self-Consistency. These findings establish a deployment hierarchy for SLMs, demonstrating that SFT combined with simplified test-time scaling is superior to complex agentic workflows for edge-based reasoning.

小模型越南语推理增强测试时扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。