用合成数据提升法语社交媒体情绪分析,降低成本与隐私风险
Model in Distress: Sentiment Analysis on French Synthetic Social Media

- 通过微调模型回译生成170万条法语合成推文,构建可扩展数据管道
- 训练6亿参数推理模型,在人工标注数据上达77-79%准确率,媲美商用大模型
- 方法适用于多语言场景,保护用户隐私且降低标注成本
社交媒体客户反馈的自动化分析面临三大挑战:标注数据成本高、多语言评估集稀缺,以及隐私问题导致数据难以共享和复现。本文通过构建通用的合成数据生成流程,解决法语公共交通客户焦虑检测的案例问题。该方法利用微调模型进行回译,从少量种子语料生成170万条合成推文,并补充合成推理轨迹。我们训练了6亿参数的多语言推理模型(支持英法双语推理),在人工标注的测试集上达到77%-79%的准确率,性能与当前最优的专有大模型及专用编码器相当。该方法不仅显著降低标注成本,还通过避免真实用户数据暴露来保障隐私。其范式可推广至其他任务与语言。
原文摘要 · Abstract (English)
Automated analysis of customer feedback on social media is hindered by three challenges: the high cost of annotated training data, the scarcity of evaluation sets, especially in multilingual settings, and privacy concerns that prevent data sharing and reproducibility. We address these issues by developing a generalizable synthetic data generation pipeline applied to a case study on customer distress detection in French public transportation. Our approach utilizes backtranslation with fine-tuned models to generate 1.7 million synthetic tweets from a small seed corpus, complemented by synthetic reasoning traces. We train 600M-parameter reasoners with English and French reasoning that achieve 77-79% accuracy on human-annotated evaluation data, matching or exceeding SOTA proprietary LLMs and specialized encoders. Beyond reducing annotation costs, our pipeline preserves privacy by eliminating the exposure of sensitive user data. Our methodology can be adopted for other use cases and languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。