Dataset catalog
Filter by type, access, and pricing. Specs show before you open the product page.
Tulu 3 SFT Mixture
Allen AI's 939k-row supervised fine-tuning mixture combining public and synthetic instruction data used to train the Tulu 3 models.
Orca AgentInstruct 1M v1
1,046,410 synthetic instruction/response conversations generated by Microsoft's AgentInstruct agentic pipeline across 15 task-type splits.
Gymnasium (Farama)
The standard Python RL API and reference environment suite (successor to OpenAI Gym) covering Classic Control, Box2D, Toy Text, MuJoCo, and Atari.
OpenMathInstruct-2
14M math problem-solution pairs (GSM8K/MATH augmentation) generated by Llama-3.1-405B-Instruct; text chain-of-thought solutions.
SWE-bench / SWE-bench Verified
Benchmark of real GitHub-issue tasks (2,294 full; 500 human-verified) evaluated by FAIL_TO_PASS/PASS_TO_PASS unit tests in Docker.
AppWorld
Simulated world of 9 apps and 457 APIs with 750 interactive coding tasks evaluated by state-based unit tests.
VisualWebArena
Multimodal, visually grounded web-agent benchmark: 910 tasks over self-hosted Classifieds, Shopping and Reddit sites.
WorkArena / WorkArena++ (ServiceNow)
Enterprise knowledge-work benchmark on live ServiceNow instances: 33 L1 atomic tasks (19,912 instances) plus 682 L2/L3 compositional tasks.
Infinity-Instruct
BAAI's large-scale open SFT collection; multi-config (7M chat, 3M foundational, plus dated Gen sets). Gated on Hugging Face.
tau-bench (τ-bench)
Tool-agent-user benchmark of 165 customer-service tasks (retail 115, airline 50) with policy-following and pass@k evaluation.
WildChat-1M
Real user–ChatGPT (GPT-3.5/GPT-4) conversations collected by Ai2 with metadata and moderation labels; current train split ~838K conversations.
WebArena
Realistic self-hosted web environment: 812 long-horizon tasks over shopping, forum, GitLab, CMS and maps with execution-based evaluation.
CodeFeedback-Filtered-Instruction
156.5K high-quality single-turn code instructions filtered (complexity 4-5 via Qwen-72B-Chat) from four open code-instruction sources; M-A-P.
orca-math-word-problems-200k
200k grade-school math word problems with GPT-4-Turbo-generated worked solutions; English, text explanations only.
Magicoder OSS-Instruct 75K & Evol-Instruct 110K
ISE-UIUC's paired Magicoder code-instruction datasets: OSS-Instruct (75K, seeded from open-source snippets) and Evol-Instruct (110K, decontaminated evol-codealpaca).
OpenHermes 2.5
Teknium's ~1M-row compilation of primarily GPT-4-generated instruction, chat, coding and reasoning data in ShareGPT (from/value) format.
Nectar (7-wise Ranked Preference Data)
182,954 chat prompts each with 7 ranked responses (GPT-4-judged) — the preference dataset behind Berkeley's Starling reward model.
SlimOrca
Open-Orca's ~518k-row curated subset of OpenOrca — GPT-4 FLAN reasoning traces with GPT-4 verification against human annotations.
UltraChat 200k
HuggingFaceH4's heavily filtered subset of UltraChat (~208k train_sft dialogues) used to train the Zephyr chat models.
MetaMathQA
395k math QA pairs bootstrapped from GSM8K and MATH via rephrasing, self-verification and backward (FOBAR) augmentation; text solutions only.