Dataset catalog
Filter by type, access, and pricing. Specs show before you open the product page.
Terminal-Bench
Benchmark of 89 hard, realistic command-line tasks run in Docker; agents issue tmux/bash keystrokes verified by outcome tests.
AgentGym (14 environments)
Unified framework of 14 interactive environments across 7 scenario types with a common HTTP/ReAct interface, plus trajectory datasets and the AgentEval benchmark.
tau2-bench (τ²-bench)
Dual-control tool-agent benchmark (278 tasks: retail 114, telecom 114, airline 50) where both agent and user can call tools.
LIBERO
Human-teleoperated demonstration data for the LIBERO lifelong robot-manipulation benchmark: 6,500 successful trajectories (50 per task) across 130 language-conditioned tasks in four suites, stored as robomimic-format HDF5.
RAGEN Environments
Reinforcement-learning framework with 10 stylized interactive environments (Sokoban, FrozenLake, Bandit, Countdown, Sudoku, WebShop, etc.) for training multi-turn reasoning agents.
TextArena
Open collection of 100+ competitive/cooperative text games with an OpenAI-Gym-style interface, online play, and a TrueSkill leaderboard for LLM agents.
R2E-Gym
Procedurally-curated executable SWE gym of 8,135+ Dockerized Python bug-fix environments with unit tests for training agents.
xLAM Function-Calling 60k
60,000 verifiable single/multi/parallel function-calling instances (query + available tools + reference tool-call answers) generated and triple-verified by Salesforce's APIGen pipeline.
BrowserGym
Unified Gymnasium harness for web agents aggregating MiniWoB, WebArena, VisualWebArena, WorkArena, AssistantBench, WebLINX and more.
SWE-Gym
Executable RL environment of 2,438 real Python GitHub-issue tasks with runtimes and unit tests for training SWE agents.
Gorilla APIBench
Instruction-to-API-call dataset over HuggingFace, TorchHub, and TensorHub APIs, used to train and evaluate LLMs that write correct API calls (the Gorilla / APIBench benchmark).
OSWorld
Real-computer benchmark: 369 open-ended desktop/web tasks on Ubuntu (also Windows/macOS) with execution-based, script evaluation.
Orca AgentInstruct 1M v1
1,046,410 synthetic instruction/response conversations generated by Microsoft's AgentInstruct agentic pipeline across 15 task-type splits.
ToolACE
11,300 synthetic multi-turn function-calling dialogues generated over a self-evolved pool of 26,507 APIs and filtered by a dual-layer (rule-based + model-based) verification pipeline.
SWE-bench / SWE-bench Verified
Benchmark of real GitHub-issue tasks (2,294 full; 500 human-verified) evaluated by FAIL_TO_PASS/PASS_TO_PASS unit tests in Docker.
AppWorld
Simulated world of 9 apps and 457 APIs with 750 interactive coding tasks evaluated by state-based unit tests.
VisualWebArena
Multimodal, visually grounded web-agent benchmark: 910 tasks over self-hosted Classifieds, Shopping and Reddit sites.
WorkArena / WorkArena++ (ServiceNow)
Enterprise knowledge-work benchmark on live ServiceNow instances: 33 L1 atomic tasks (19,912 instances) plus 682 L2/L3 compositional tasks.
tau-bench (τ-bench)
Tool-agent-user benchmark of 165 customer-service tasks (retail 115, airline 50) with policy-following and pass@k evaluation.
WebArena
Realistic self-hosted web environment: 812 long-horizon tasks over shopping, forum, GitLab, CMS and maps with execution-based evaluation.