arXiv:2608.18097cs.CL2026-08

构建首个跨媒体法语新闻编辑部分类基准,提升多来源新闻内容理解能力。

FrenchNews-7: Benchmarking Cross-Publisher French News Editorial Desk Classification

论文配图:FrenchNews-7: Benchmarking Cross-Publisher French News Editorial Desk Classification
图 1 · 摘自论文原文
  • 融合网址信息与大模型标注,构建7类法语新闻编辑部标签体系
  • 细调CamemBERT模型在全文章输入下召回率达0.799,优于零样本大模型
  • 揭示经济与社会类栏目边界模糊,适合关注法语媒体分析的研究者

我们提出FrenchNews-7,一个基于法国多媒体来源的法语新闻编辑部分类基准,包含大规模多渠道语料、基于网址的七类分类体系,以及微调后的CamemBERT分类器。标签通过结合出版商网址片段与大模型标注的混合流程生成,并经双人双模型一致性检验(成对κ≥0.766,人-人间κ=0.806)。在分布内与跨出版商设置下评估词法、多语言及法语特化训练模型,同时对比三个零样本大模型(GPT-OSS-120B、Mistral Small 3.2、Llama-3.3-70B)在未见出版商数据上的表现。最优模型CamemBERT-base在完整文本输入下整体召回率达0.799,显著优于所有零样本大模型,差距集中于经济与社会两类模糊类别。跨出版商评估显示边界稳定性不均:体育、文化与国际类转移清晰,而经济类召回仅0.517,接近人工盲评一致率(0.55),社会类精度0.577,反映其边界模糊性源于编辑惯例而非模型能力不足。相关模型、数据集与标注脚本已公开于Hugging Face。

原文摘要 · Abstract (English)

We present FrenchNews-7, a cross-publisher France-based French-language news editorial desk classification benchmark combining a large multi-outlet corpus, a URL-derived seven-class taxonomy, and a fine-tuned CamemBERT classifier. Labels are assigned via a hybrid pipeline combining publisher URL slugs with LLM annotation for structurally ambiguous cases, audited through an inter-rater study (2 humans + 2 LLMs; pairwise $κ\geq 0.766$, human--human $κ= 0.806$). We evaluate lexical, multilingual, and French-specific trained classifiers under both in-distribution and held-out-publisher settings, with additional comparison against zero-shot LLM baselines (GPT-OSS-120B, Mistral Small 3.2, Llama-3.3-70B) on the held-out pool. The strongest model, CamemBERT-base on full article text, outperforms headline-only input, generalizes to unseen outlets, and exceeds all three zero-shot LLM baselines on overall recall (0.799), with the gap concentrated in the ambiguous editorial-boundary categories Economie and Societe. Cross-publisher evaluation reveals uneven boundary stability: Sport, Culture & Loisirs, and International transfer cleanly, while Economie (recall = 0.517) is close to blinded human agreement (0.55), and Societe (precision = 0.577) absorbs boundary ambiguity, both suggesting editorial conventions rather than recoverable classifier headroom. The fine-tuned CamemBERT-base model, labeled manifest, reference collection scripts, and a reliability-tier guidance table are available at https://huggingface.co/LeFrenchNewsLab/camembert-base-frenchnews7 (model) and https://huggingface.co/datasets/LeFrenchNewsLab/frenchnews-7 (dataset).

新闻分类法语NLPCamemBERT多源数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。