GENERAL AGENT BENCHMARKS · 2026

通用 Agent Bench
从长轨迹到持续演化

这里不把“会反思的方法”误写成“新 benchmark”。每项工作都按任务载体、时间尺度、反馈来源、可保留记忆、最终 evaluator 和原始公开入口拆开。

单 episode 长轨迹OSWorld · WorkArena++ · AgencyBench · SWE-bench · τ-bench
→
反馈驱动迭代MLAgentBench · OPT-BENCH · DySQL
→
跨 episode 学习LLM-Evolve · LifelongAgentBench · LIBERO · AgentRecBench
→
能力资产化Reflexion · ExpeL · Agent-Pro · MetaReflection · Mem²Evolve
01 / LANDSCAPE

先区分:数据集、评测框架还是方法

场景很长不等于会演化;会重试也不等于跨任务学习。卡片上的类型与时间尺度决定它能回答什么研究问题。

02 / INSTANCE AUDIT

OSWorld:官方任务与 evaluator

NeurIPS 2024 · Datasets & Benchmarks · 独立 Benchmark

I have calculated the total work hours from the everyday hours... multiply the total hours with the hourly rate... Help me fill in the cell the correct answer. Don't touch irrelevant blank regions.
  1. snapshot: libreoffice_calc
  2. source: Multiply_Time_Number.xlsx
  3. target artifact: /home/user/Multiply_Time_Number.xlsx
  4. evaluator: compare_table · Sheet 0 · E3

审计说明:英文任务为论文或官方仓库可核对的原文 / 最小节选;中文内容解释构造与评分逻辑,不冒充官方数据。请通过上方“原始 Demo”链接逐字段 double-check。