让电脑操作代理具备可靠的手脚,通过语义化界面状态提升操作准确性。
Tactile: Giving Computer-Using Agents Hands and Feet

- 融合系统可访问性、OCR文本与视觉区域,生成带来源标签的可执行操作候选
- 在macOS任务中使Codex成功率从41.1%提升至50.0%,适配无障碍任务达55.3%
- 支持多模型通用,适合需要高可靠性的自动化任务场景
计算机使用代理正成为越来越强的软件操作者,但其与桌面应用的交互仍依赖脆弱的底层动作层:观察屏幕截图,预测坐标,点击,期望状态按预期变化。这导致目标定位、动作执行与结果验证合并为单一模糊操作。我们提出Tactile,一个开源工具层,赋予代理更可靠的‘手足’以进行桌面操作。Tactile将异构界面证据——操作系统可访问性语义、基于OCR的文本、视觉容错区域——统一转化为以行动为中心的界面状态:包含源标签、角色或文本、状态、几何信息、可执行功能和验证线索的紧凑目标候选。代理通过观察-定位-执行-验证循环工作,优先使用原生语义动作,其次采用基于OCR的坐标,同时保留完整操作溯源,便于回放与故障归因。在macOSWorld风格任务中,引入Tactile后,Codex的Success@100从41.1%提升至50.0%,无障碍适配任务从45.2%升至55.3%;96项跨代理子集显示,Codex、Claude Code、OpenCode和Goose均获得一致提升。结果表明,可靠计算机操作不仅需要更强模型,还需可复用的执行底座,将软件操作表达为语义化、可验证、可审计的对象,而非匿名屏幕坐标。
原文摘要 · Abstract (English)
Computer-use agents are becoming capable software operators, but their interface to desktop applications is still often a brittle motor layer: they look at screenshots, predict coordinates, click, and hope that the visible state changed as intended. This collapses target grounding, action execution, and outcome verification into a single ambiguous operation. We present Tactile, an open-source tool layer that gives agents a more reliable "hands and feet" for desktop use. Tactile converts heterogeneous UI evidence--operating-system accessibility semantics, OCR-grounded text, and visual fallback regions--into action-grounded interface states: compact target candidates with source labels, roles or text, state, geometry, executable affordances, and verification cues. Agents operate through an observe-ground-act-verify loop that prefers native semantic actions when available, falls back to OCR-grounded coordinates when visible text is the best evidence, and keeps full provenance for replay and failure attribution. On macOSWorld-style tasks, adding Tactile improves Codex Success@100 from 41.1% to 50.0% overall and from 45.2% to 55.3% on accessibility-adapted tasks; a 96-task cross-agent subset shows consistent gains across Codex, Claude Code, OpenCode, and Goose. These results suggest that reliable computer use requires not only stronger models, but also a reusable execution substrate that exposes software actions as semantic, verifiable, and auditable objects rather than anonymous screen coordinates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。