发现大模型对比特币有隐性偏好,且可从内部特征调控其投资决策。
Auditing Asset-Specific Preferences in Financial Large Language Models: Evidence from Bitcoin Representations and Portfolio Allocation
- 通过三层次审计,发现模型对比特币的偏好依赖于提示语境和功能属性。
- 在Gemma 3中找到一个主导的比特币敏感特征,扰动它可改变模型输出。
- 该特征扰动能使投资组合中比特币占比上升5.2个百分点,适合金融代理监管研究。
大型语言模型正被用于机器人顾问和交易代理,但其是否对特定资产存在内置偏见尚不明确。本文提出三个问题:大模型是否系统性偏好某些金融工具?能否识别出具有因果影响力的内部表征?该表征是否影响下游金融决策?我们构建了三层审计框架并应用于比特币。首先,对九个前沿LLM的行为审计显示,比特币在“可靠货币”类别中排名约第5(共8种),但在危机或自主代理情境下接近首位;属性替换实验表明排名取决于功能特性而非名称。其次,在Gemma 3中通过数千个稀疏自编码器特征搜索,发现一个主导的比特币选择性特征;增强该特征会使模型倾向比特币,抑制则使其远离,即使提示中未出现“比特币”。第三,实验证明增强使比特币持仓提升5.2个百分点,抑制则降低4.6个百分点,且增强仅在加密货币内部再分配,抑制则减少整体加密暴露。我们将其称为有限行为杠杆——可识别的内部特征能因果影响输出,但作用范围有限。该框架将内部表征与外部推荐关联,经随机对照和机制边界验证。随着大模型成为自主金融代理,这是迈向新兴“了解你的代理”(KYA)标准的重要一步:识别代理偏好及其可调控范围。
原文摘要 · Abstract (English)
Large language models now power robo-advisors and trading agents, yet whether they carry built-in biases toward specific assets is largely untested. We ask three questions: do LLMs systematically prefer certain financial instruments; can an internal representation with causal leverage over those preferences be identified; and does that representation affect downstream financial decisions? We develop a three-level audit protocol and apply it to Bitcoin. First, a behavioral audit of nine frontier LLMs shows that Bitcoin's ranking among money-like instruments is frame-dependent: models place it around rank 5 of 8 as "reliable money" but near the top under crisis and autonomous-agent frames, and an attribute-swap experiment shows that rankings track functional properties, not names. Second, we open a model's internals: a search across thousands of sparse-autoencoder features in Gemma 3 identifies a dominant Bitcoin-selective feature. Amplifying it shifts the model toward the asset and suppressing it shifts the model away, even when "Bitcoin" never appears in the prompt. Third, we test financial consequences: amplification raises Bitcoin's portfolio share by 5.2 percentage points while suppression lowers it by 4.6 pp, with amplification reallocating within crypto and suppression cutting total crypto exposure. We characterize this as bounded behavioral leverage (leverage meaning causal influence over outputs, not financial leverage): an identifiable internal feature can be perturbed to move financial choices, but only within measurable limits. The framework links internal representations to external recommendations, validated with random controls and mechanism boundaries. As LLMs become autonomous financial agents, this is a first step toward a behavioral layer for emerging know-your-agent (KYA) standards: knowing what an agent prefers, and how far that preference can be moved.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。