Convex Markets / Datasets / All departments

Dataset catalog

Filter by type, access, and pricing. Specs show before you open the product page.

32 results for “Reasoning”
M
External · RQA-5009

MMLU-Pro

A harder, reasoning-focused successor to MMLU: 12,032 multiple-choice questions across 14 subjects with up to 10 options each, scored by exact match on the gold answer letter.

Type Reasoning QAVolume 12032 multiple-choice questions (test split; plus 70 validation CoT few-shot examples)Format ParquetAccess PUBLIC LICENSE
Free
open dataset (HF, MIT)
View
G
External · RQA-5004

GPQA

GPQA is a gated benchmark of 448 expert-written, 'Google-proof' graduate-level multiple-choice questions in biology, physics, and chemistry, built for reasoning evaluation and scalable-oversight research.

Type Reasoning QAVolume 448 multiple-choice questions (GPQA Main set; also GPQA Diamond=198, GPQA Extended=546)Format CSVAccess GATED
Free
open dataset (HF, gated)
View
RG
External · RLE-1403

Reasoning Gym

Python library of 100+ procedural dataset generators with algorithmic verifiers for RL with verifiable rewards; generates virtually unlimited reasoning problems with adjustable difficulty.

Type RL environmentsVolume 100+ generatorsFormat Python (pip: reasoning-gym)Access PUBLIC LICENSE
Free
open-source license
View
N
External · SFT-2310

Nemotron-Post-Training-Dataset-v1

NVIDIA 25.6M-row post-training corpus (chat/code/math/stem/tool_calling) — includes a rare 310K genuine tool-calling split with real tool_calls schemas; CC-BY-4.0.

Type SFT datasetVolume 25659642 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
BH
External · RQA-5008

BIG-Bench Hard (BBH)

A 6,511-example reasoning benchmark of 23 hard BIG-Bench tasks (input/target pairs) used to test chain-of-thought prompting, graded by exact match on the gold answer.

Type Reasoning QAVolume 6,511 examples (test split; across 27 task configs / 23 BBH tasks)Format parquet (Hugging Face); per-task JSON in the source GitHub repoAccess PUBLIC LICENSE
Free
open dataset (HF)
View
O
External · SFT-2307

OpenThoughts-114k

114K verified DeepSeek-R1 reasoning traces over math, science, code and puzzles; Open Thoughts, Apache-2.0.

Type SFT datasetVolume 114000 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
S
External · RLE-1402

Search-R1

Open-source RL framework that trains LLMs to interleave reasoning with live search-engine calls; the retriever is treated as part of the RL environment.

Type RL environmentsVolume retrieval-augmented QA training frameworkFormat Python (built on veRL; installed from source via conda/pip)Access PUBLIC LICENSE
Free
open-source license
View
O
External · SFT-2205

OpenMathReasoning

~5.68M math solutions over 306k unique AoPS problems, split into CoT, tool-integrated reasoning (Python code) and GenSelect.

Type SFT datasetVolume ~5.68M solutionsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
RE
External · RLE-1401

RAGEN Environments

Reinforcement-learning framework with 10 stylized interactive environments (Sokoban, FrozenLake, Bandit, Countdown, Sudoku, WebShop, etc.) for training multi-turn reasoning agents.

Type RL environmentsVolume 10 environmentsFormat Python (Gym-compatible interface; installed from source via setup script)Access PUBLIC LICENSE
Free
open-source license
View
T
External · RLE-1306

TextArena

Open collection of 100+ competitive/cooperative text games with an OpenAI-Gym-style interface, online play, and a TrueSkill leaderboard for LLM agents.

Type RL environmentsVolume 100+ gamesFormat Python (pip: textarena) + Gym-style APIAccess PUBLIC LICENSE
Free
open-source license
View
L(
External · SFT-2309

Llama-Nemotron-Post-Training-Dataset (v1.1)

NVIDIA post-training corpus for Llama-Nemotron models spanning math, code, science, instruction-following, chat and safety; CC-BY-4.0.

Type SFT datasetVolume 33011757 samplesFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
O(
External · SFT-2305

OpenCodeReasoning (OCR-1)

735K competitive-programming reasoning samples (Python) with R1-generated chain-of-thought over 28,319 unique questions; NVIDIA, CC-BY-4.0.

Type SFT datasetVolume 735255 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
B
External · SFT-2308

Bespoke-Stratos-17k

16.7K reasoning traces (~10K math, ~5K code, ~1K science/puzzle) distilled from DeepSeek-R1 via the Sky-T1 pipeline; Bespoke Labs, Apache-2.0.

Type SFT datasetVolume 16710 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
M(
External · RQA-5002

MATH (Hendrycks)

12,500 competition mathematics problems (7,500 train / 5,000 test) with full step-by-step LaTeX solutions whose final answer is wrapped in \boxed{}.

Type Reasoning QAVolume 12,500 problems (7,500 train / 5,000 test)Format Parquet (Hugging Face mirror, 7 subject configs); original repository ships per-problem JSON files with fields problem/level/type/solutionAccess PUBLIC LICENSE
Free
open dataset (HF)
View
OA
External · SFT-2105

Orca AgentInstruct 1M v1

1,046,410 synthetic instruction/response conversations generated by Microsoft's AgentInstruct agentic pipeline across 15 task-type splits.

Type SFT datasetVolume 1046410 examplesFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
O
External · SFT-2204

OpenMathInstruct-2

14M math problem-solution pairs (GSM8K/MATH augmentation) generated by Llama-3.1-405B-Instruct; text chain-of-thought solutions.

Type SFT datasetVolume 14M problem-solution pairsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
W/
External · RLE-1104

WorkArena / WorkArena++ (ServiceNow)

Enterprise knowledge-work benchmark on live ServiceNow instances: 33 L1 atomic tasks (19,912 instances) plus 682 L2/L3 compositional tasks.

Type RL environmentsVolume 33 L1 atomic tasks (19,912 instances) + 682 WorkArena++ L2/L3 tasksFormat Python package (pip install browsergym-workarena) + live ServiceNow instance + PlaywrightAccess GATED
Free
open-source license
View
o
External · SFT-2206

orca-math-word-problems-200k

200k grade-school math word problems with GPT-4-Turbo-generated worked solutions; English, text explanations only.

Type SFT datasetVolume 200k word-problem QA pairsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
AA
External · RQA-5003

AI2 ARC

A dataset of 7,787 genuine grade-school-level multiple-choice science questions, split into a harder Challenge Set and an Easy Set, for evaluating advanced question answering and reasoning.

Type Reasoning QAVolume 7787 questionsFormat parquetAccess PUBLIC LICENSE
Free
open dataset (HF)
View
O2
External · SFT-2002

OpenHermes 2.5

Teknium's ~1M-row compilation of primarily GPT-4-generated instruction, chat, coding and reasoning data in ShareGPT (from/value) format.

Type SFT datasetVolume 1001551 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View