低资源语言法律理解中,模型推理能力是失败主因,非语料问题。
Cross Lingual Transfer in Tulu Legal Comprehension: Script-Dependent Improvement and RAG-Induced Knowledge Conflict
- 通过转写德鲁维达语系文字提升模型初步理解能力
- 检索式问答引发事实替换与虚构错误,表现依赖脚本类型
- 揭示模型自身推理缺陷,适用于多语言低资源评估
在低资源语言环境中,缺乏充足训练语料时,常借助高资源语言作为辅助。以法律领域为背景,测试了Llama3、Hex-1和Sarvam三个模型对低资源德拉威语(图鲁语)法律诉状的分类能力。通过跨德拉威语系脚本的转写,模型无需大规模训练即可获得初步理解,但理解程度高度依赖脚本类型(卡纳达语——另一低资源语言——表现最优)。在RAG框架下从卡纳达语法律文献库中检索信息后,部分模型在特定条件下表现出微弱正向提升,但失败时通常呈现双重问题:事实替换(过度聚焦片段导致推理偏移)与幻觉(无依据的编造)。研究发现,在低资源场景中,推理失败主要源于模型对信息的解析与后续推断能力,而非语料本身。脚本依赖性与RAG鲁棒性密切相关,这一结论经推理轨迹分析与统计诚实性框架验证,具有更广泛适用性。
原文摘要 · Abstract (English)
Low-resource languages without an adequate training corpus often use a related, higher-resource language as a scaffold for comprehension. Still, there is a need to develop rigorous evaluation methods to identify when models fail in cross lingual low-resource environments. Using the legal domain as a backdrop, three models (Llama3, Hex-1, Sarvam) were tested on the ability to classify legal complaints written in a low resource Dravidian language (Tulu). Transliterating queries across Dravidian scripts allowed models to gain a preliminary understanding of speakers' complaints without the use of wide scale training, though the level of comprehension was heavily script dependent (with Kannada - another relatively low-resource language - producing the strongest positive trend). Retrieving from a corpus of Kannada legal papers across a RAG framework caused mixed results. Some models had a weak positive trend in comprehension under certain conditions, but when models failed, it was often across two axes: fact substitution (fixating on specific passage excerpts that skewed reasoning) and confabulation (hallucination that had no basis in either query or corpus). Within low resource domains, results identify the model's parsing of information and subsequent reasoning as the source of reasoning failure, rather than corpus contents. Script-dependent comprehension and RAG robustness also seem to travel together. This is further supported by the reasoning-trace analysis and a statistical-honesty framework deployed - techniques that are more broadly applicable to low-resource multilingual RAG evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。