Convex Markets / Datasets / SFT dataset

Dataset catalog

Filter by type, access, and pricing. Specs show before you open the product page.

39 results
N
External · SFT-2310

Nemotron-Post-Training-Dataset-v1

NVIDIA 25.6M-row post-training corpus (chat/code/math/stem/tool_calling) — includes a rare 310K genuine tool-calling split with real tool_calls schemas; CC-BY-4.0.

Type SFT datasetVolume 25659642 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
O
External · SFT-2307

OpenThoughts-114k

114K verified DeepSeek-R1 reasoning traces over math, science, code and puzzles; Open Thoughts, Apache-2.0.

Type SFT datasetVolume 114000 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
O
External · SFT-2205

OpenMathReasoning

~5.68M math solutions over 306k unique AoPS problems, split into CoT, tool-integrated reasoning (Python code) and GenSelect.

Type SFT datasetVolume ~5.68M solutionsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
L(
External · SFT-2309

Llama-Nemotron-Post-Training-Dataset (v1.1)

NVIDIA post-training corpus for Llama-Nemotron models spanning math, code, science, instruction-following, chat and safety; CC-BY-4.0.

Type SFT datasetVolume 33011757 samplesFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
O(
External · SFT-2305

OpenCodeReasoning (OCR-1)

735K competitive-programming reasoning samples (Python) with R1-generated chain-of-thought over 28,319 unique questions; NVIDIA, CC-BY-4.0.

Type SFT datasetVolume 735255 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
S
External · SFT-2003

SmolTalk

Hugging Face TB's ~1.04M-row synthetic SFT mixture (the 'all' config) used to train the SmolLM2 instruct models.

Type SFT datasetVolume 1043917 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
B
External · SFT-2308

Bespoke-Stratos-17k

16.7K reasoning traces (~10K math, ~5K code, ~1K science/puzzle) distilled from DeepSeek-R1 via the Sky-T1 pipeline; Bespoke Labs, Apache-2.0.

Type SFT datasetVolume 16710 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
T3
External · SFT-2001

Tulu 3 SFT Mixture

Allen AI's 939k-row supervised fine-tuning mixture combining public and synthetic instruction data used to train the Tulu 3 models.

Type SFT datasetVolume 939344 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
OA
External · SFT-2105

Orca AgentInstruct 1M v1

1,046,410 synthetic instruction/response conversations generated by Microsoft's AgentInstruct agentic pipeline across 15 task-type splits.

Type SFT datasetVolume 1046410 examplesFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
O
External · SFT-2204

OpenMathInstruct-2

14M math problem-solution pairs (GSM8K/MATH augmentation) generated by Llama-3.1-405B-Instruct; text chain-of-thought solutions.

Type SFT datasetVolume 14M problem-solution pairsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
I
External · SFT-2104

Infinity-Instruct

BAAI's large-scale open SFT collection; multi-config (7M chat, 3M foundational, plus dated Gen sets). Gated on Hugging Face.

Type SFT datasetVolume 7449106 examplesFormat Parquet (Hugging Face)Access GATED
Free
open-source license
View
W
External · SFT-2107

WildChat-1M

Real user–ChatGPT (GPT-3.5/GPT-4) conversations collected by Ai2 with metadata and moderation labels; current train split ~838K conversations.

Type SFT datasetVolume 837989 conversationsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
NR
External · SFT-2404

No Robots

10,000 human-written instruction-and-demonstration pairs across 10 task categories, created by skilled human annotators (no model-generated data) for supervised fine-tuning.

Type SFT datasetVolume 10,000 instruction-demonstration pairsFormat parquetAccess PUBLIC LICENSE
Free
open dataset (HF)
View
C
External · SFT-2306

CodeFeedback-Filtered-Instruction

156.5K high-quality single-turn code instructions filtered (complexity 4-5 via Qwen-72B-Chat) from four open code-instruction sources; M-A-P.

Type SFT datasetVolume 156526 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
o
External · SFT-2206

orca-math-word-problems-200k

200k grade-school math word problems with GPT-4-Turbo-generated worked solutions; English, text explanations only.

Type SFT datasetVolume 200k word-problem QA pairsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
AD
External · SFT-2407

Aya Dataset

204,112 human-authored multilingual instruction prompt-completion pairs across 65 languages, curated by native/fluent speakers via Cohere Labs' Aya Annotation Platform for multilingual instruction tuning.

Type SFT datasetVolume 204,112 prompt-completion pairs (202,362 train + 1,750 test)Format ParquetAccess PUBLIC LICENSE
Free
open dataset (HF)
View
OC
External · SFT-2406

OpenAssistant Conversations v2 (OASST2)

A human-generated, human-annotated corpus of multilingual assistant-style conversation trees released by the OpenAssistant community for supervised fine-tuning and alignment research.

Type SFT datasetVolume 135,174 messages (128,575 train + 6,599 validation)Format ParquetAccess PUBLIC LICENSE
Free
open dataset (HF)
View
MO
External · SFT-2301

Magicoder OSS-Instruct 75K & Evol-Instruct 110K

ISE-UIUC's paired Magicoder code-instruction datasets: OSS-Instruct (75K, seeded from open-source snippets) and Evol-Instruct (110K, decontaminated evol-codealpaca).

Type SFT datasetVolume 186380 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
S
External · SFT-2005

SlimOrca

Open-Orca's ~518k-row curated subset of OpenOrca — GPT-4 FLAN reasoning traces with GPT-4 verification against human annotations.

Type SFT datasetVolume 517982 rowsFormat JSONLAccess PUBLIC LICENSE
Free
open-source license
View
O2
External · SFT-2002

OpenHermes 2.5

Teknium's ~1M-row compilation of primarily GPT-4-generated instruction, chat, coding and reasoning data in ShareGPT (from/value) format.

Type SFT datasetVolume 1001551 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View