Dataset catalog
Filter by type, access, and pricing. Specs show before you open the product page.
Skywork-Reward-Preference-80K
77,016 curated chosen/rejected chat preference pairs subsampled from public sources (HelpSteer2, OffsetBias, WildGuard, Magpie DPO series) and used to train the Skywork-Reward reward models.
HelpSteer2
NVIDIA's open (CC-BY-4.0) helpfulness dataset of 21,362 human-annotated prompt-response samples rated on five attributes, with a dedicated human pairwise-preference split for training reward models and DPO.
PKU-SafeRLHF
A large human-annotated safety-preference dataset where each question has two model responses ranked separately for helpfulness and harmlessness, plus per-response safety meta-labels across 19 harm categories.
UltraFeedback
A large-scale, fine-grained preference dataset of ~64k prompts, each with 4 model completions rated by GPT-4 across four aspects, for training reward and critique models.
Intel Orca DPO Pairs
A ~12.9K-example preference dataset in Direct Preference Optimization (DPO) format, derived from Open-Orca/OpenOrca, pairing a 'chosen' and 'rejected' response for each instruction prompt.
Nectar (7-wise Ranked Preference Data)
182,954 chat prompts each with 7 ranked responses (GPT-4-judged) — the preference dataset behind Berkeley's Starling reward model.
Chatbot Arena Conversations
33K real pairwise human preference votes from the LMSYS Chatbot Arena, where users compared two anonymous LLMs side by side and chose the better response.
Stanford Human Preferences (SHP)
385,563 collective human preference pairs mined from 18 Reddit Q&A subreddits, where the more-helpful of two comments is inferred from Reddit scores and timestamps.
WebGPT Comparisons
19,578 human-scored answer-pair comparisons from OpenAI's WebGPT project, used to train the WebGPT long-form question-answering reward model.
Anthropic HH-RLHF
Anthropic's human preference dataset of chosen/rejected assistant dialogue pairs for training helpful and harmless (HH) reward models via RLHF.