用英语思考,韩语回答,高效适配多语言智能体。
Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents

- 基于英语指令训练模型,输出韩语响应,提升跨语言效率。
- 在单卡部署下实现4比特量化,推理速度提升3倍以上。
- 适用于企业级多语言智能体,支持复杂工具调用任务。
我们提出LuckyStar 111B,一个1110亿参数的混合推理模型,由Cohere与LG CNS合作开发,专为韩英双语企业级智能体设计,满足实际内存与服务限制。该模型基于Cohere的完全后训练命令模型(Command A)进行微调,而非重新预训练,并采用前导条件机制,在简洁非推理模式与长流程工具导向推理间切换。研究了四种高效扩展多语言工具使用智能体的方法:多语言监督微调、基于可验证奖励的强化学习(用于多步工具调用任务)、保持韩语输出语言一致性的奖励机制,以及单GPU部署下的4比特量化。适应后的模型在数学推理、函数调用及自然语言转SQL(NL2SQL)任务中表现提升,同时维持了韩语与英语指令遵循能力。结果为后训练多语言模型在资源受限环境下的可验证智能体工作流适配提供了实用方案与失败模式分析。
原文摘要 · Abstract (English)
We present LuckyStar 111B, a 111B-parameter hybrid reasoning model developed through a collaboration between Cohere and LG CNS for Korean-English enterprise agents under practical memory and serving constraints. The model trains from Cohere's fully post-trained Command A model rather than a new pretraining run, and uses preamble conditioning to switch between concise non-reasoning behavior and longer tool-oriented reasoning. We study four choices for scaling tool-using agents efficiently: multilingual supervised fine-tuning, reinforcement learning with verifiable rewards for multi-step tool-use tasks, language-consistency rewards for Korean user-facing responses, and 4-bit quantization for single-GPU serving. The adapted model improves mathematical reasoning, function calling, and agentic natural-language-to-SQL (NL2SQL) performance while preserving general Korean and English instruction-following quality. These results provide a practical recipe and failure-mode analysis for adapting post-trained multilingual models to verifiable agentic workflows under memory-constrained deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。