arXiv:2508.11278cs.HCcs.AI2025-08被引 4

AI在软件工程中也会受语言线索影响而产生认知偏差,且越复杂任务越明显。

Is General-Purpose AI Reasoning Sensitive to Data-Induced Cognitive Biases? Dynamic Benchmarking on Typical Software Engineering Dilemmas

  • 用动态生成任务测试AI是否受语言暗示误导
  • 主流AI对偏见敏感度6%-35%,复杂任务达49%
  • 适合关注AI可靠性与工程落地风险的研究者

软件工程中的认知偏差可能导致高昂错误。尽管通用人工智能(GPAI)因非人类特性可能缓解此类偏差,但其训练数据来自人类,引发关键问题:GPAI自身是否也存在认知偏差?为此,我们提出首个动态基准框架,评估GPAI在软件工程流程中由数据诱发的认知偏差。基于16个手工设计的真实任务,每个任务包含8种认知偏差(如锚定效应、框架效应)及其无偏变体,测试无关任务逻辑的语言提示能否导致GPAI从正确转向错误结论。为扩大规模并保证真实感,我们开发基于GPAI的按需增强管道,生成保留偏见提示但改变表面细节的任务变体,确保正确率88%-99%(经人工评估),提升多样性,并通过Prolog推理控制推理复杂度。评估GPT、LLaMA、DeepSeek等主流GPAI系统发现,它们普遍依赖浅层语言启发式而非深层推理。所有系统均表现出偏见敏感性(6%-35%),且随任务复杂度上升至49%,揭示了AI驱动软件工程中的潜在风险。

原文摘要 · Abstract (English)

Human cognitive biases in software engineering can lead to costly errors. While general-purpose AI (GPAI) systems may help mitigate these biases due to their non-human nature, their training on human-generated data raises a critical question: Do GPAI systems themselves exhibit cognitive biases? To investigate this, we present the first dynamic benchmarking framework to evaluate data-induced cognitive biases in GPAI within software engineering workflows. Starting with a seed set of 16 hand-crafted realistic tasks, each featuring one of 8 cognitive biases (e.g., anchoring, framing) and corresponding unbiased variants, we test whether bias-inducing linguistic cues unrelated to task logic can lead GPAI systems from correct to incorrect conclusions. To scale the benchmark and ensure realism, we develop an on-demand augmentation pipeline relying on GPAI systems to generate task variants that preserve bias-inducing cues while varying surface details. This pipeline ensures correctness (88-99% on average, according to human evaluation), promotes diversity, and controls reasoning complexity by leveraging Prolog-based reasoning. We evaluate leading GPAI systems (GPT, LLaMA, DeepSeek) and find a consistent tendency to rely on shallow linguistic heuristics over more complex reasoning. All systems exhibit bias sensitivity (6-35%), which increases with task complexity (up to 49%) and highlights risks in AI-driven software engineering.

AI偏见软件工程动态评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。