构建首个覆盖印度恒河平原多语言的文学文本问答数据集
LittiChoQA: Literary Texts in Indic Languages Chosen for Question Answering
- 基于网络收集的文学文本自动生成27万+问答对
- 全上下文微调下模型最高语义得分76.1,短上下文提速显著
- 适合低资源语言、文学理解与长文本问答研究者
针对现代大模型在文学文本上的长上下文问答难题,尤其在低资源语言中,我们提出了LittiChoQA——迄今最大的印地语系文学问答数据集,涵盖印度恒河平原多种语言。该数据集包含超过27万条自动生成的问答对,覆盖事实型与非事实型问题,来源为从公开网络获取的真实文学文本。我们在非事实型、抽象问答任务上评估多个多语言大模型,对比全上下文与上下文缩短两种设置。结果表明性能与效率存在明显权衡:全上下文微调在词级与语义级评分上表现最优,而上下文缩短显著提升吞吐量。其中Krutrim-2表现最强,在全上下文下语义得分为76.1;在上下文缩短时,采用段落选择得分为74.9,向量检索得分为71.4。定性分析进一步验证了上述结论。
原文摘要 · Abstract (English)
Long-context question answering (QA) over literary texts poses significant challenges for modern large language models, particularly in low-resource languages. We address the scarcity of long-context QA resources for Indic languages by introducing LittiChoQA, the largest literary QA dataset to date covering many languages spoken in the Gangetic plains of India. The dataset comprises over 270K automatically generated question-answer pairs with a balanced distribution of factoid and non-factoid questions, generated from naturally authored literary texts collected from the open web. We evaluate multiple multilingual LLMs on non-factoid, abstractive QA, under both full-context and context-shortened settings. Results demonstrate a clear trade-off between performance and efficiency: full-context fine-tuning yields the highest token-level and semantic-level scores, while context shortening substantially improves throughput. Among the evaluated models, Krutrim-2 achieves the strongest performance, obtaining a semantic score of 76.1 with full context. While, in shortened context settings it scores 74.9 with answer paragraph selection and 71.4 with vector-based retrieval. Qualitative evaluations further corroborate these findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。