arXiv:2605.08590cs.HCcs.AI2026-05被引 2

检测大模型生成的个人感知解释中过度推断的问题

Causal Stories from Sensor Traces: Auditing Epistemic Overreach in LLM-Generated Personal Sensing Explanations

论文配图:Causal Stories from Sensor Traces: Auditing Epistemic Overreach in LLM-Generated Personal Sensing Explanations
图 1 · 摘自论文原文
  • 用结构化标准评估大模型解释的证据支撑度
  • 发现模型在95%以上场景中存在无据推断
  • 适合关注可解释性与可信推理的研究者

大语言模型被用于解释个人感知数据,将活动与情绪轨迹转化为异常日发生原因的自然语言描述。然而,这些解释可能看似连贯且富有意义,即使证据稀少或缺失。本文提出‘认知过度推断’(Epistemic Overreach, EO)作为衡量标准,用于识别生成解释超出可用传感证据支持的情况。我们从三个纵向感知数据集(StudentLife、GLOBEM、CollegeExperience)中获取异常日场景,涵盖活动、睡眠和情绪异常,使用Llama、Qwen、GPT三类模型,在两种提示条件下生成14,922条解释:一种为最小约束提示,另一种明确要求模型仅基于数据作限域陈述。通过结构化评分表,将EO分解为五维:无支持的因果归因、未承认的数据缺口、过度自信语言、时间不一致性和诊断性推断。结果表明,模型普遍存在无依据归因,该现象在不同数据集、异常类型和模型家族间均重复出现。增加行为证据并未显著降低EO;限域提示有所缓解但无法消除问题。研究建议将证据根基性作为评估个人感知解释的首要标准,超越流畅性和合理性。个人感知解释需建立证据纪律:系统必须区分观察事实、推断内容与未知信息。

原文摘要 · Abstract (English)

LLMs are increasingly used to explain personal sensing data, translating traces of activity and mood into natural-language accounts of why an anomalous day may have occurred. However, such explanations can sound coherent and personally meaningful even when the underlying evidence is sparse or missing. We introduce epistemic overreach (EO) as a measure for cases where a generated explanation implies more than the available sensing evidence can justify. To audit how often and in what forms EO occurs, we obtained anomalous-day scenarios from three longitudinal sensing datasets of college students: StudentLife, GLOBEM, and CollegeExperience. Across activity, sleep, and affect anomalies, we generated 14,922 explanations using three LLM families -- Llama, Qwen, and GPT -- under two prompting conditions: one minimally constrained prompt and another prompt explicitly instructing models to bound claims to the data. For each scenario, we varied the amount of behavioral evidence available to the model to examine whether more evidence reduces EO. We evaluated each explanation using a structured rubric, decomposing EO into the dimensions of unsupported causal attribution, unacknowledged data gaps, overconfident language, temporal inconsistency, and diagnostic inference. We find that LLMs routinely attribute anomalous days to causes without sufficient support from the data, and that this pattern replicates across datasets, anomaly types, and model families. Further, providing richer context does not reliably reduce EO; bounded prompting helps but does not eliminate it. These findings suggest that evidential grounding should be a first-order evaluation criterion for LLM-generated personal sensing explanations, alongside fluency and plausibility. We argue that personal sensing explanations require evidential discipline: systems must distinguish what is observed, what is inferred, and what remains unknown.

大模型解释认知偏差证据验证个人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。