SKU SFT-2309 · Sold by External
Llama-Nemotron-Post-Training-Dataset (v1.1)
Product specifications
| SKU | SFT-2309 |
|---|---|
| Data type | SFT dataset |
| Volume | 33011757 samples |
| Size on disk | 130 GB |
| Format | Parquet (Hugging Face) |
| Access model | PUBLIC LICENSE |
| Pricing | Free · open-source license |
| Quality score | — |
| License | CC-BY-4.0 (some subsets ODC-BY / CC-BY-SA per card) |
NVIDIA's post-training dataset used to build the Llama-Nemotron reasoning models (v1.1). The card's per-category sample counts are: math 22,066,397; code 10,108,883; science 708,920; instruction following 56,339; chat 39,792; safety 31,426. Configs/splits: an SFT config (code, math, science, chat, safety) and an RL config (instruction_following). Responses were synthetically generated by DeepSeek-R1, Qwen-2.5 and Llama-3.x models. No dedicated tool/function-calling split — it is reasoning, code, chat and safety data only.