Convex Markets / Datasets / SFT dataset / Nemotron-Post-Training-Dataset-v1
SKU SFT-2310 · Sold by External

Nemotron-Post-Training-Dataset-v1

Product specifications

SKUSFT-2310
Data typeSFT dataset
Volume25659642 rows
Size on disk203 GB
FormatParquet (Hugging Face)
Access modelPUBLIC LICENSE
PricingFree · open-source license
Quality score
LicenseCC-BY-4.0
NVIDIA's Nemotron-Post-Training-Dataset-v1 is a 25,659,642-row post-training corpus with per-split counts: chat 746,622; code 1,896,395; math 2,044,407; stem 20,662,167; tool_calling 310,051. Synthetic responses were generated by DeepSeek-R1-0528 and Qwen3-235B-A22B. Its standout feature is the tool_calling split (~310K rows): unlike the code-interpreter reasoning found elsewhere, these rows contain genuine structured function-calling traces — a tools specification plus assistant messages carrying tool_calls, covering single-turn, multi-turn and multi-step tool use. This is the rare authentic function-calling SFT data in this catalog.