Dataset catalog
Filter by type, access, and pricing. Specs show before you open the product page.
SlimOrca
Open-Orca's ~518k-row curated subset of OpenOrca — GPT-4 FLAN reasoning traces with GPT-4 verification against human annotations.
Nectar (7-wise Ranked Preference Data)
182,954 chat prompts each with 7 ranked responses (GPT-4-judged) — the preference dataset behind Berkeley's Starling reward model.
AgentInstruct (AgentTuning)
1,866 reward-filtered, ReAct-style multi-turn agent interaction trajectories (GPT-4 generated) spanning six real-world agent tasks, released to fine-tune the AgentLM models from the AgentTuning paper.
Language-Table
A large suite of human-collected, language-conditioned tabletop block-manipulation demonstrations from Robotics at Google, where a robot pushes colored blocks in response to natural-language instructions, released in RLDS/TFDS format.
ToolBench (ToolLLM)
126,486 instruction-tuning instances pairing real user queries over 16,464 real-world RapidAPI APIs with ChatGPT-generated DFSDT (depth-first search decision tree) tool-use solution paths and final answers.
UltraChat 200k
HuggingFaceH4's heavily filtered subset of UltraChat (~208k train_sft dialogues) used to train the Zephyr chat models.
Glaive Function Calling v2
112,960 synthetic multi-turn chat conversations pairing a SYSTEM tool-schema prompt with assistant function-call turns, released by Glaive AI on Hugging Face under Apache-2.0.
BridgeData V2
60,096 real WidowX 250 robot manipulation trajectories (teleoperated + scripted) across 24 environments and 13 skills, each labeled with natural-language task instructions.
MetaMathQA
395k math QA pairs bootstrapped from GSM8K and MATH via rephrasing, self-verification and backward (FOBAR) augmentation; text solutions only.
MathInstruct
262k math instruction examples combining chain-of-thought and program-of-thought (executable Python) rationales compiled from 13 datasets.
Evol-CodeAlpaca v1
111K English code instruction/output pairs — an Evol-Instruct augmentation of CodeAlpaca-20k using GPT-4 across 10 evolution strategies.
Chatbot Arena Conversations
33K real pairwise human preference votes from the LMSYS Chatbot Arena, where users compared two anonymous LLMs side by side and chose the better response.
Dolphin
Open Orca-style FLAN reproduction: ~892K FLANv2 completions from GPT-4 plus ~2.84M from GPT-3.5, filtered to remove refusals/alignment.
WizardLM Evol-Instruct V2 196k
Complexity-evolved (Evol-Instruct) instruction conversations derived from Alpaca/ShareGPT for WizardLM SFT; HF split ships 143K rows.
OpenOrca
An open collection of ~2.94M FLAN-Collection instructions paired with GPT-4/GPT-3.5 'Orca-style' reasoning-trace responses for supervised instruction fine-tuning.
Evol-Instruct-Code-80k-v1
78K code instruction/output pairs — an open reproduction of WizardCoder's Evol-Instruct-Code (CodeAlpaca run through 3 evolution rounds).
ToolAlpaca
A corpus of ~3,900 generalized tool-use agent trajectories over 400+ real-world APIs, generated by a multi-agent (user / assistant / tool-executor) simulation and formatted as ReAct-style Thought/Action/Action Input/Observation traces.
Self-Instruct
Machine-generated instruction-following dataset created by bootstrapping GPT-3 from 175 seed tasks, released as ~82K prompt/completion demonstrations for instruction tuning.
LIMA
GAIR's 1,030-example gated dataset of high-quality curated prompts and responses used in the 'Less Is More for Alignment' study.
ALOHA / ACT Demonstrations
Human-teleoperated bimanual fine-manipulation demonstrations collected with the low-cost open-source ALOHA hardware for the ACT paper, recorded as multi-camera RGB video plus 14-DoF joint states/actions in per-episode HDF5 files.