为什么重要

该研究揭示了AI代理在时间感知方面的局限性,对依赖AI进行任务规划和资源管理的场景具有重要启示,有助于改进代理设计。

关键事实

事实 1

研究测试了 Anthropic 的 Claude Code 和 OpenAI 的 Codex 这两款编码助手的时间感。

来源与依据

单一来源

The pair tested two widely used coding assistants, Anthropic's Claude Code and OpenAI's Codex, on their sense of time.

THE DECODER · 第一方证据 · 支持

查看 THE DECODER 原文

事实 2

在 ProgramBench 测试中,两款模型大多预估需要约 90 分钟,与任务难度无关。

来源与依据

单一来源

On ProgramBench, both models mostly guessed around 90 minutes, no matter the difficulty.

THE DECODER · 第一方证据 · 支持

查看 THE DECODER 原文

事实 3

当代理获得报告已用时间的工具时,它们几乎每次都能准确判断时间。

来源与依据

单一来源

When the agents got access to a tool that reports elapsed time, they got it right almost every time.

THE DECODER · 第一方证据 · 支持

查看 THE DECODER 原文