模型的训练方法比所属家族更能决定其对话行为差异。
Post-Training Recipe, More Than Model Family, Shapes Multi-Agent LLM Conversational Behavior
- 通过对比同基础模型不同训练版本的对话表现,发现训练配方影响更大。
- 同一Llama模型在不同搭档下语气保守度变化达18%,超过跨家族差异。
- 适合关注多智能体系统设计与模型选择的研究者参考。
多语言模型系统利用多个大模型进行推理、互评或协作。其价值取决于相同输入下各模型是否产生可测量的行为差异。以往离线研究建议每家族选一个模型以保证多样性,因大模型在孤立评估时倾向于偏好同家族输出。但这一家族标签在真实交互式多模型系统中的有效性尚未验证。本文基于94万条链的11个检查点语料和160万条同基础Llama因子实验数据,发现,在经验证的“避险”(hedging)指标上,同一Llama检查点的回应行为随搭档不同变动达18%,超过任何跨家族差距。Qwen、闭源API及运行时测试支持该模式非孤立现象,修复与挑战分析尚存局限因表面线索检测器能力不足。结果表明,后训练配方是多模型组合的首要维度,仅凭模型家族无法充分预测对话多样性。
原文摘要 · Abstract (English)
Multi-LLM systems use multiple language models to deliberate, judge each other's outputs, or coordinate as agents. Their value depends on the models producing measurably different conversational behaviors when given the same input. Prior offline studies recommend drawing one model per family for behavioral diversity, because LLMs prefer outputs from their own family when rating one another in isolation. Whether the same family label predicts behavior in interactive multi-LLM systems, the setting that real deployed systems use, has not been tested. We study this with a 940,000-chain 11-checkpoint corpus and a 1.6M-chain same-base Llama factorial. On our validated headline metric, hedging, a reasoning-distilled Llama checkpoint shifts by 18% depending on which same-base partner it replies to, more than any cross-family hedging gap in the controlled subset. Qwen, closed-API, and runtime checks suggest the pattern is not isolated, while repair and challenge analyses remain exploratory because their surface-cue detectors are weaker. Overall, the results identify post-training recipe as a first-class axis for multi-LLM panel composition and show that model family alone is an incomplete proxy for conversational diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。