Dataset catalog
Filter by type, access, and pricing. Specs show before you open the product page.
Reasoning Gym
Python library of 100+ procedural dataset generators with algorithmic verifiers for RL with verifiable rewards; generates virtually unlimited reasoning problems with adjustable difficulty.
Nemotron-Post-Training-Dataset-v1
NVIDIA 25.6M-row post-training corpus (chat/code/math/stem/tool_calling) — includes a rare 310K genuine tool-calling split with real tool_calls schemas; CC-BY-4.0.
OpenThoughts-114k
114K verified DeepSeek-R1 reasoning traces over math, science, code and puzzles; Open Thoughts, Apache-2.0.
Search-R1
Open-source RL framework that trains LLMs to interleave reasoning with live search-engine calls; the retriever is treated as part of the RL environment.
OpenMathReasoning
~5.68M math solutions over 306k unique AoPS problems, split into CoT, tool-integrated reasoning (Python code) and GenSelect.
RAGEN Environments
Reinforcement-learning framework with 10 stylized interactive environments (Sokoban, FrozenLake, Bandit, Countdown, Sudoku, WebShop, etc.) for training multi-turn reasoning agents.
TextArena
Open collection of 100+ competitive/cooperative text games with an OpenAI-Gym-style interface, online play, and a TrueSkill leaderboard for LLM agents.
Llama-Nemotron-Post-Training-Dataset (v1.1)
NVIDIA post-training corpus for Llama-Nemotron models spanning math, code, science, instruction-following, chat and safety; CC-BY-4.0.
OpenCodeReasoning (OCR-1)
735K competitive-programming reasoning samples (Python) with R1-generated chain-of-thought over 28,319 unique questions; NVIDIA, CC-BY-4.0.
Bespoke-Stratos-17k
16.7K reasoning traces (~10K math, ~5K code, ~1K science/puzzle) distilled from DeepSeek-R1 via the Sky-T1 pipeline; Bespoke Labs, Apache-2.0.
Orca AgentInstruct 1M v1
1,046,410 synthetic instruction/response conversations generated by Microsoft's AgentInstruct agentic pipeline across 15 task-type splits.
OpenMathInstruct-2
14M math problem-solution pairs (GSM8K/MATH augmentation) generated by Llama-3.1-405B-Instruct; text chain-of-thought solutions.
WorkArena / WorkArena++ (ServiceNow)
Enterprise knowledge-work benchmark on live ServiceNow instances: 33 L1 atomic tasks (19,912 instances) plus 682 L2/L3 compositional tasks.
orca-math-word-problems-200k
200k grade-school math word problems with GPT-4-Turbo-generated worked solutions; English, text explanations only.
SlimOrca
Open-Orca's ~518k-row curated subset of OpenOrca — GPT-4 FLAN reasoning traces with GPT-4 verification against human annotations.
OpenHermes 2.5
Teknium's ~1M-row compilation of primarily GPT-4-generated instruction, chat, coding and reasoning data in ShareGPT (from/value) format.
MetaMathQA
395k math QA pairs bootstrapped from GSM8K and MATH via rephrasing, self-verification and backward (FOBAR) augmentation; text solutions only.
MathInstruct
262k math instruction examples combining chain-of-thought and program-of-thought (executable Python) rationales compiled from 13 datasets.
WizardLM Evol-Instruct V2 196k
Complexity-evolved (Evol-Instruct) instruction conversations derived from Alpaca/ShareGPT for WizardLM SFT; HF split ships 143K rows.
OpenR1-Math-220k
220k competition math problems, each with 2-4 DeepSeek-R1 chain-of-thought traces verified by Math-Verify and Llama-3.3-70B; text reasoning, no code.