arXiv:2607.27611cs.CLq-fin.RM2026-07

用AI精准提取企业外汇对冲披露信息,可审计且结果可信。

AWARE-FX: An Auditable Knowledge-Guided AI System for Measuring Corporate Foreign-Exchange Hedging Disclosure

  • 结合专业词典与逻辑规则,从年报中自动识别对冲证据。
  • 在超2.4万家企业年份数据中提取54万条文本片段,准确率达0.872。
  • 适合金融监管、财报分析及需要可追溯AI决策的研究者使用。

企业年报中关于外汇风险管理、衍生品使用、自然对冲及明确不使用的信息呈弱结构化特征。本研究构建了AWARE-FX系统,一种可审计的AI/NLP决策支持工具,将报告文本转化为可追踪的企业年度对冲披露度量。系统融合专业来源词典、否定与会计状态逻辑、渠道特定财务编码器、精确证据门控、保守聚合策略及审计日志。在2008-2025年间覆盖24,909个香港企业年份,成功检索并评分543,527条文本片段。通过消融实验、分层300条样本人工审计、三种种子模型(FinBERT-ModernBERT)对比、严格2023-2025年时间测试、概率校准、选择性预测及固定提示生成模型基准,验证其可靠性。FinBERT在八项任务中的七项平均F1更高,时间序列F1范围为0.702至0.872;对置信度最低20%样本弃权后,保留样本F1提升0.050至0.077。确定性Qwen3-8B在商品和否定证据上表现良好,但在外债与会计语境标签上表现差,说明通用大模型无法完全替代领域约束。严格外汇得分与关联基线及压力期外汇暴露负相关,而通用宽泛得分无此关联,提供外部构念效度支持,非对冲有效性因果估计。AWARE-FX贡献了一个经验证的决策支持架构,其中检索、状态逻辑、分类、不确定性处理、聚合与外部验证均保持独立可审计。

原文摘要 · Abstract (English)

Corporate annual reports contain weakly structured evidence about foreign-exchange risk management, derivative use, natural hedging, and explicit non-use. This study develops AWARE-FX, an auditable AI/NLP decision-support system that converts report text into traceable firm-year hedging-disclosure measures. The system combines a professional-source lexicon, negation and accounting-status logic, channel-specific financial encoders, exact evidence gates, conservative aggregation, and an audit ledger. Across 24,909 Hong Kong firm-years from 2008-2025, it retrieves and scores 543,527 snippets. Reliability is evaluated through ablations, a stratified 300-snippet human audit, three-seed FinBERT-ModernBERT comparisons, strict 2023-2025 temporal tests, probability calibration, selective prediction, and fixed-prompt generative-model benchmarks. FinBERT has the higher mean F1 in seven of eight encoder task-split comparisons; its temporal F1 ranges from 0.702 to 0.872. Abstaining on the 20% least-confident temporal observations raises retained-sample F1 by 0.050-0.077. Deterministic Qwen3-8B performs strongly on commodity and negation evidence but poorly on foreign-debt and accounting-context labels, showing that a general-purpose LLM does not uniformly replace domain constraints. The strict FX score is negatively associated with linked baseline and stress-period FX exposure, whereas the generic broad score is not. These associations provide external construct validation, not causal estimates of hedging effectiveness. AWARE-FX contributes a tested decision-support architecture in which retrieval, status logic, classification, uncertainty handling, aggregation, and external validation remain separately auditable.

AI审计财报分析NLP应用外汇对冲

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。