arXiv:2502.12932cs.CL2025-02被引 2

用大模型生成爪哇语和巽他语的有文化背景的故事,提升低资源语言推理能力。

Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese

  • 用文化提示引导大模型生成故事,替代昂贵的人工标注。
  • 基于大模型生成的数据训练后,下游任务性能优于翻译或印尼语数据。
  • 适合研究低资源语言、文化敏感性推理及生成式数据构建的学者。

由于数据稀缺和本地标注成本高昂,低资源语言中的文化根基常识推理尚未得到充分探索。我们测试大型语言模型(LLMs)在该场景下生成具有文化细腻度叙事的可行性。聚焦爪哇语和巽他语,比较三种数据生成策略:(1) 使用文化线索提示的大模型辅助生成故事,(2) 从印尼语基准数据集机器翻译,(3) 本地人撰写的故事。人工评估显示,大模型生成的故事在文化契合度上接近母语者,但在连贯性和正确性上仍有差距。在每种数据集上微调模型,并在人工编写的测试集上评估分类与生成任务。结果显示,大模型生成的数据在下游任务中表现优于机器翻译数据和印尼语人工数据。本文发布一个高质量的爪哇语和巽他语文化根基常识故事基准数据集,以支持后续研究。

原文摘要 · Abstract (English)

Culturally grounded commonsense reasoning is underexplored in low-resource languages due to scarce data and costly native annotation. We test whether large language models (LLMs) can generate culturally nuanced narratives for such settings. Focusing on Javanese and Sundanese, we compare three data creation strategies: (1) LLM-assisted stories prompted with cultural cues, (2) machine translation from Indonesian benchmarks, and (3) native-written stories. Human evaluation finds LLM stories match natives on cultural fidelity but lag in coherence and correctness. We fine-tune models on each dataset and evaluate on a human-authored test set for classification and generation. LLM-generated data yields higher downstream performance than machine-translated and Indonesian human-authored training data. We release a high-quality benchmark of culturally grounded commonsense stories in Javanese and Sundanese to support future work.

低资源语言文化推理生成数据多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。