Convex Markets / Datasets / All departments

Dataset catalog

Filter by type, access, and pricing. Specs show before you open the product page.

110 results
PI
External · RLE-1405

Prime Intellect Environments Hub / verifiers

Community hub and Python registry of open-source RL environments, built on the MIT-licensed verifiers library; site headlines 2,500+ environments as of Jul 2026.

Type RL environmentsVolume 2500+ environmentsFormat Python wheels via Prime CLI (uv / pyproject); verifiers library (MIT)Access PUBLIC LICENSE
Free
open-source license
View
C
External · RLE-1501

Craftax

Craftax is a lightning-fast, JAX-native open-ended RL benchmark that reimplements and extends Crafter with NetHack-inspired roguelike mechanics, using ~67 unlockable achievements as sparse-reward tasks.

Type RL environmentsVolume 67 achievements (sparse-reward tasks in full Craftax; Craftax-Classic has 22)Format Python package (JAX / gymnax interface); pip install craftaxAccess PUBLIC LICENSE
Free
open-source license (MIT)
View
M
External · RQA-5009

MMLU-Pro

A harder, reasoning-focused successor to MMLU: 12,032 multiple-choice questions across 14 subjects with up to 10 options each, scored by exact match on the gold answer letter.

Type Reasoning QAVolume 12032 multiple-choice questions (test split; plus 70 validation CoT few-shot examples)Format ParquetAccess PUBLIC LICENSE
Free
open dataset (HF, MIT)
View
G
External · RQA-5004

GPQA

GPQA is a gated benchmark of 448 expert-written, 'Google-proof' graduate-level multiple-choice questions in biology, physics, and chemistry, built for reasoning evaluation and scalable-oversight research.

Type Reasoning QAVolume 448 multiple-choice questions (GPQA Main set; also GPQA Diamond=198, GPQA Extended=546)Format CSVAccess GATED
Free
open dataset (HF, gated)
View
T
External · RLE-1204

Terminal-Bench

Benchmark of 89 hard, realistic command-line tasks run in Docker; agents issue tmux/bash keystrokes verified by outcome tests.

Type RL environmentsVolume 89 tasksFormat Docker + YAML/dir task specs + Python harnessAccess PUBLIC LICENSE
Free
open-source license
View
RG
External · RLE-1403

Reasoning Gym

Python library of 100+ procedural dataset generators with algorithmic verifiers for RL with verifiable rewards; generates virtually unlimited reasoning problems with adjustable difficulty.

Type RL environmentsVolume 100+ generatorsFormat Python (pip: reasoning-gym)Access PUBLIC LICENSE
Free
open-source license
View
S
External · RLE-1404

SkyRL-Gym

Gymnasium-API library of tool-use environments (math, code, search, text-to-SQL) for LLM post-training, part of the SkyRL RL stack from NovaSky.

Type RL environmentsVolume tool-use environment library (math/code/search/SQL)Format Python (Gymnasium API; pip)Access PUBLIC LICENSE
Free
open-source license
View
A(
External · RLE-1305

AgentGym (14 environments)

Unified framework of 14 interactive environments across 7 scenario types with a common HTTP/ReAct interface, plus trajectory datasets and the AgentEval benchmark.

Type RL environmentsVolume 14 environmentsFormat Python + HTTP env servers (agentenv)Access PUBLIC LICENSE
Free
open-source license
View
N
External · SFT-2310

Nemotron-Post-Training-Dataset-v1

NVIDIA 25.6M-row post-training corpus (chat/code/math/stem/tool_calling) — includes a rare 310K genuine tool-calling split with real tool_calls schemas; CC-BY-4.0.

Type SFT datasetVolume 25659642 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
BH
External · RQA-5008

BIG-Bench Hard (BBH)

A 6,511-example reasoning benchmark of 23 hard BIG-Bench tasks (input/target pairs) used to test chain-of-thought prompting, graded by exact match on the gold answer.

Type Reasoning QAVolume 6,511 examples (test split; across 27 task configs / 23 BBH tasks)Format parquet (Hugging Face); per-task JSON in the source GitHub repoAccess PUBLIC LICENSE
Free
open dataset (HF)
View
t(
External · RLE-1207

tau2-bench (τ²-bench)

Dual-control tool-agent benchmark (278 tasks: retail 114, telecom 114, airline 50) where both agent and user can call tools.

Type RL environmentsVolume 278 tasks (retail 114 + telecom 114 + airline 50)Format Python package + JSON domain dataAccess PUBLIC LICENSE
Free
open-source license
View
O
External · SFT-2307

OpenThoughts-114k

114K verified DeepSeek-R1 reasoning traces over math, science, code and puzzles; Open Thoughts, Apache-2.0.

Type SFT datasetVolume 114000 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
S
External · RLE-1402

Search-R1

Open-source RL framework that trains LLMs to interleave reasoning with live search-engine calls; the retriever is treated as part of the RL environment.

Type RL environmentsVolume retrieval-augmented QA training frameworkFormat Python (built on veRL; installed from source via conda/pip)Access PUBLIC LICENSE
Free
open-source license
View
L
External · DEM-6005

LIBERO

Human-teleoperated demonstration data for the LIBERO lifelong robot-manipulation benchmark: 6,500 successful trajectories (50 per task) across 130 language-conditioned tasks in four suites, stored as robomimic-format HDF5.

Type DemonstrationsVolume 6500 human-teleoperated demonstrations (130 tasks x 50 demos)Format HDF5 (robomimic-format demonstrations) with BDDL task-definition filesAccess PUBLIC LICENSE
Free
open dataset (CC BY 4.0)
View
M
External · RLE-1506

Meta-World

Open-source benchmark of 50 simulated Sawyer-arm robotic manipulation tasks (Gymnasium API) with a shared 4-D continuous action space, dense shaped rewards, and a per-task binary success metric for multi-task and meta-RL.

Type RL environmentsVolume 50 manipulation tasks (Sawyer simulation environments; MT50 superset, MT10/MT25/ML10/ML45 subsets)Format MuJoCo simulation environments (XML models + Python env classes) exposed through the Gymnasium API; distributed as the 'metaworld' PyPI packageAccess PUBLIC LICENSE
Free
open-source license (MIT)
View
OX
External · DEM-6001

Open X-Embodiment

Open X-Embodiment is an aggregated robot-learning dataset that pools over 1 million real robot demonstration trajectories from 22 embodiments across 60 datasets into a single standardized RLDS/TFDS format.

Type DemonstrationsVolume 1,000,000+ real robot trajectories (RLDS episodes)Format RLDS episode format stored as TFDS / TFRecord (loadable via tensorflow_datasets)Access PUBLIC LICENSE
Free
open dataset (RLDS/TFDS, public GCS)
View
O
External · SFT-2205

OpenMathReasoning

~5.68M math solutions over 306k unique AoPS problems, split into CoT, tool-integrated reasoning (Python code) and GenSelect.

Type SFT datasetVolume ~5.68M solutionsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
D
External · DEM-6003

DROID

DROID is a large-scale in-the-wild robot manipulation dataset of 76,000 human-teleoperated Franka Panda demonstration trajectories (350 hours) with multi-view stereo RGB, robot state/action, and natural-language task instructions.

Type DemonstrationsVolume 76,000 teleoperated demonstration trajectories (episodes)Format RLDS / TensorFlow Datasets (tfds.load("droid")); also raw HDF5 and LeRobot parquet+mp4 conversionsAccess PUBLIC LICENSE
Free
open dataset (CC-BY 4.0)
View
RE
External · RLE-1401

RAGEN Environments

Reinforcement-learning framework with 10 stylized interactive environments (Sokoban, FrozenLake, Bandit, Countdown, Sudoku, WebShop, etc.) for training multi-turn reasoning agents.

Type RL environmentsVolume 10 environmentsFormat Python (Gym-compatible interface; installed from source via setup script)Access PUBLIC LICENSE
Free
open-source license
View
T
External · RLE-1306

TextArena

Open collection of 100+ competitive/cooperative text games with an OpenAI-Gym-style interface, online play, and a TrueSkill leaderboard for LLM agents.

Type RL environmentsVolume 100+ gamesFormat Python (pip: textarena) + Gym-style APIAccess PUBLIC LICENSE
Free
open-source license
View