arXiv:2609.01810cs.CL2026-09

首个面向波斯语对话生成与理解的统一基准,涵盖三类高质量对话数据。

TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding

论文配图:TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding
图 1 · 摘自论文原文
  • 构建三类波斯语对话数据集,经母语者多轮审核确保质量。
  • 微调仅需25-50%数据即可恢复90%以上性能提升,LoRA效果显著。
  • 适合作为波斯语NLP研究者的评估基准,尤其适合对话任务方向。

波斯语(超过1.2亿人使用)缺乏全面的对话生成与理解基准。我们推出TALKFA,一个包含三个互补数据集的统一基准:(1) WIKI-FADIAL,4.2K 基于维基百科的常识对话;(2) DAILYDIALOG-FA,6.6K 标注了对话行为与情绪的日常对话;(3) PLAYDIAL-FA,2.1K 带情感标签的戏剧对话。尽管大模型辅助构建,每条对话均经母语者多阶段审校,仅保留最终人工核准数据。六款LLAMA与MISTRAL模型实验显示,LoRA可将训练数据减少25-50%仍恢复超90%性能增益。分类任务中,FABERT在对话行为识别上表现最佳,LORA-MISTRAL-7B在情绪识别中领先,而MISTRAL-24B在情感分析中得分最高。人类评估与外部验证证实基准可靠性;与GPT-4.1对比表明,自动指标严重高估对话质量。零样本测试显示前沿大模型仍难以突破该基准。所有数据、标注指南、代码与模型检查点将公开发布。

原文摘要 · Abstract (English)

Farsi, spoken by more than 120 million people, lacks a comprehensive benchmark for dialogue generation and understanding. We introduce TALKFA, a unified benchmark comprising three complementary datasets: (1) WIKI-FADIAL, 4.2K Wikipedia-grounded dialogues for knowledge-grounded generation; (2) DAILYDIALOG-FA, 6.6K dialogues annotated for dialogue acts and emotions; and (3) PLAYDIAL-FA, 2.1K theatrical dialogues with sentiment labels. While LLMs assist data construction, every dialogue undergoes multi-stage review and revision by native Farsi speakers, and only the final human-approved dialogues are released. Experiments with six LLAMA and MISTRAL models show that LoRA substantially improves dialogue generation while requiring only 25-50% of the training data to recover over 90% of the final performance gains. Across classification tasks, FABERT achieves the best dialogue-act performance, LORA-MISTRAL-7B performs best on emotion recognition, and MISTRAL-24B achieves the highest sentiment score. Human evaluation and independent external validation demonstrate the reliability of the benchmark, while comparisons with GPT-4.1 as an LLM judge reveal that automatic metrics substantially overestimate dialogue quality. Zero-shot evaluation with frontier LLMs further shows that TalkFa remains a challenging benchmark. We will release all datasets, annotation guidelines, code, and checkpoints.

对话生成波斯语基准测试LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。