Palmyra x6通过精细化微调提升企业级智能体工具使用能力。
Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning

- 用626条真实合成轨迹,低学习率单轮微调,保留基础模型结构。
- 在BFCL核心测试中得分0.785,六项基准均值领先同行。
- 适合需要高安全性和工具调用能力的企业应用开发。
Palmyra x6 是一个专为面向企业场景的智能体任务优化的大语言模型。该模型基于混合专家(Mixture-of-Experts)基础模型,通过锚定监督微调(Anchored Supervised Fine-Tuning)在一套精炼且经验证的合成工具使用轨迹数据上进行后训练,采用 Muon + Adam 混合优化器。训练策略刻意保守:仅使用626条轨迹、单个训练轮次、低学习率,并以KL散度约束冻结的基础模型。模型在写作智能体任务中表现显著优于此前默认模型,在多个公开基准测试中表现优异,于BFCL Core测试中取得0.785分,是同组模型中六项基准平均分最高的。此外,在偏见与安全性评估中,其表现亦优于或媲美对比模型。
原文摘要 · Abstract (English)
Palmyra x6 is a large language model optimized for use with enterprise-oriented agentic tasks. The model was built by post-training a Mixture-of-Experts base model with Anchored Supervised Fine-Tuning on a compact corpus of verified, synthetic tool-use trajectories, optimized with a Muon + Adam hybrid. The recipe is deliberately conservative and deliberately controlled: 626 trajectories, a single epoch, a low learning rate, and a KL anchor to the frozen base. The model shows substantial gains over the previous default model for Writer Agent, and compares favorably with several recent models on public benchmarks, scoring the highest on BFCL Core at $0.785$ and posts the highest six-benchmark mean of the cohort. Furthermore, the model has shown itself to be competitive or leading relative to comparators in our bias and safety evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。