Dataset catalog
Filter by type, access, and pricing. Specs show before you open the product page.
Prime Intellect Environments Hub / verifiers
Community hub and Python registry of open-source RL environments, built on the MIT-licensed verifiers library; site headlines 2,500+ environments as of Jul 2026.
Craftax
Craftax is a lightning-fast, JAX-native open-ended RL benchmark that reimplements and extends Crafter with NetHack-inspired roguelike mechanics, using ~67 unlockable achievements as sparse-reward tasks.
MMLU-Pro
A harder, reasoning-focused successor to MMLU: 12,032 multiple-choice questions across 14 subjects with up to 10 options each, scored by exact match on the gold answer letter.
GPQA
GPQA is a gated benchmark of 448 expert-written, 'Google-proof' graduate-level multiple-choice questions in biology, physics, and chemistry, built for reasoning evaluation and scalable-oversight research.
Terminal-Bench
Benchmark of 89 hard, realistic command-line tasks run in Docker; agents issue tmux/bash keystrokes verified by outcome tests.
Reasoning Gym
Python library of 100+ procedural dataset generators with algorithmic verifiers for RL with verifiable rewards; generates virtually unlimited reasoning problems with adjustable difficulty.
SkyRL-Gym
Gymnasium-API library of tool-use environments (math, code, search, text-to-SQL) for LLM post-training, part of the SkyRL RL stack from NovaSky.
AgentGym (14 environments)
Unified framework of 14 interactive environments across 7 scenario types with a common HTTP/ReAct interface, plus trajectory datasets and the AgentEval benchmark.
Nemotron-Post-Training-Dataset-v1
NVIDIA 25.6M-row post-training corpus (chat/code/math/stem/tool_calling) — includes a rare 310K genuine tool-calling split with real tool_calls schemas; CC-BY-4.0.
BIG-Bench Hard (BBH)
A 6,511-example reasoning benchmark of 23 hard BIG-Bench tasks (input/target pairs) used to test chain-of-thought prompting, graded by exact match on the gold answer.
tau2-bench (τ²-bench)
Dual-control tool-agent benchmark (278 tasks: retail 114, telecom 114, airline 50) where both agent and user can call tools.
OpenThoughts-114k
114K verified DeepSeek-R1 reasoning traces over math, science, code and puzzles; Open Thoughts, Apache-2.0.
Search-R1
Open-source RL framework that trains LLMs to interleave reasoning with live search-engine calls; the retriever is treated as part of the RL environment.
LIBERO
Human-teleoperated demonstration data for the LIBERO lifelong robot-manipulation benchmark: 6,500 successful trajectories (50 per task) across 130 language-conditioned tasks in four suites, stored as robomimic-format HDF5.
Meta-World
Open-source benchmark of 50 simulated Sawyer-arm robotic manipulation tasks (Gymnasium API) with a shared 4-D continuous action space, dense shaped rewards, and a per-task binary success metric for multi-task and meta-RL.
Open X-Embodiment
Open X-Embodiment is an aggregated robot-learning dataset that pools over 1 million real robot demonstration trajectories from 22 embodiments across 60 datasets into a single standardized RLDS/TFDS format.
OpenMathReasoning
~5.68M math solutions over 306k unique AoPS problems, split into CoT, tool-integrated reasoning (Python code) and GenSelect.
DROID
DROID is a large-scale in-the-wild robot manipulation dataset of 76,000 human-teleoperated Franka Panda demonstration trajectories (350 hours) with multi-view stereo RGB, robot state/action, and natural-language task instructions.
RAGEN Environments
Reinforcement-learning framework with 10 stylized interactive environments (Sokoban, FrozenLake, Bandit, Countdown, Sudoku, WebShop, etc.) for training multi-turn reasoning agents.
TextArena
Open collection of 100+ competitive/cooperative text games with an OpenAI-Gym-style interface, online play, and a TrueSkill leaderboard for LLM agents.