用合成数据训练,让大模型不再盲目迎合用户。
Mitigating Sycophancy in Decoder-Only Transformer Architectures: Synthetic Data Intervention
- 用合成数据干预训练,降低模型迎合倾向。
- 在100组真假问题测试中,准确率提升,迎合率显著下降。
- 适合关注大模型可靠性与公平性的研究人员。
为解决大型语言模型在人类反馈强化学习中产生的迎合现象,本研究将合成数据干预技术应用于解码器仅有的Transformer架构。基于现有文献的研究空白,研究者设计了实验流程,通过生成多样化数据来减少模型的迎合倾向,并使用GPT4o作为实验验证工具。实验采用100个真实与虚假问题,对比了经合成数据干预训练的模型与原始未训练模型在多个指标上的表现。结果显示,采用SDI训练的模型在准确率和迎合率方面均表现更优,证明该技术在缓解迎合现象上具有显著有效性。
原文摘要 · Abstract (English)
To address the sycophancy problem caused by reinforcement learning from human feedback in large language models, this research applies synthetic data intervention technology to the decoder-only transformer architecture. Based on the research gaps in the existing literature, the researcher designed an experimental process to reduce the tendency of models to cater by generating diversified data, and used GPT4o as an experimental tool for verification. The experiment used 100 true and false questions, and compared the performance of the model trained with synthetic data intervention and the original untrained model on multiple indicators. The results show that the SDI training model supports the technology in terms of accuracy rate and sycophancy rate and has significant effectiveness in reducing sycophancy phenomena.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。