研究土耳其语证据形态如何受信源可信度影响,发现人类能敏感识别,但大模型表现差。
Benchmarking Source-Sensitive Reasoning in Turkish: Humans and LLMs under Evidential Trust Manipulation

- 通过操控信源可信度,测试土耳其语过去时态后缀-DI与-mIs的选择机制。
- 母语者在高可信度情境下更多用-DI,低可信度则倾向-mIs,结果稳定。
- 10个大模型多数表现不一致或反转,显示人类与模型在信源推理上存在明显差距。
本文研究信源可信度是否影响土耳其语证据形态,以及大语言模型(LLMs)能否捕捉这种敏感性。在控制的填空任务中,信息来源明确为外部,仅其可信度被操纵(高可信度 vs. 低可信度)。人类生产实验显示,母语者表现出显著的可信度效应:高可信度情境中更倾向于使用后缀-DI,低可信度情境中则更多使用-mIs,该模式在多种敏感性分析中保持稳定。随后评估了10个LLMs在三种提示范式下的表现(开放式填空、显式过去时填空、强制选择A/B)。LLM行为高度依赖模型和提示方式:部分模型出现微弱或局部一致的可信度响应,但整体效果不稳定,常出现反转,并频繁被输出合规问题及强烈的基础后缀偏好所掩盖。结果为基于信任/承诺的土耳其语证据性理论提供新证据,揭示了人类与大模型在源敏感证据推理上的显著差距。
原文摘要 · Abstract (English)
This paper investigates whether source trustworthiness shapes Turkish evidential morphology and whether large language models (LLMs) track this sensitivity. We study the past-domain contrast between -DI and -mIs in controlled cloze contexts where the information source is overtly external, while only its perceived reliability is manipulated (High-Trust vs. Low-Trust). In a human production experiment, native speakers of Turkish show a robust trust effect: High-Trust contexts yield relatively more -DI, whereas Low-Trust contexts yield relatively more -mIs, with the pattern remaining stable across sensitivity analyses. We then evaluate 10 LLMs in three prompting paradigms (open gap-fill, explicit past-tense gap-fill, and forced-choice A/B selection). LLM behavior is highly model- and prompt-dependent: some models show weak or local trust-consistent shifts, but effects are generally unstable, often reversed, and frequently overshadowed by output-compliance problems and strong base-rate suffix preferences. The results provide new evidence for a trust-/commitment-based account of Turkish evidentiality and reveal a clear human-LLM gap in source-sensitive evidential reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。