SKU PRF-3001 · Sold by External

UltraFeedback

Product specifications

SKUPRF-3001
Data typePreference data
Volume63967 prompts (train split rows; 256k completions total)
Size on disk940 MB
FormatParquet (HF datasets); JSON in source repo
Access modelPUBLIC LICENSE
PricingFree · open dataset (HF)
Quality score
LicenseMIT License
UltraFeedback is a large-scale, fine-grained, diverse preference dataset built by OpenBMB (THUNLP) to train reward and critique models for RLHF. About 63,967 instructions were sampled from six public sources (UltraChat, ShareGPT, Evol-Instruct, TruthfulQA, FalseQA, and FLAN) and used to query a pool of 17 LLMs, sampling 4 responses per prompt for 256k completions in total. GPT-4 then annotates each completion under a fine-grained instruction covering four aspects (instruction-following, truthfulness, honesty, and helpfulness), producing both numerical ratings and textual rationales plus an overall_score. The card notes RLHF researchers can construct roughly 340k comparison pairs from these ratings. A December 2023 update corrected an overall_score bug that had mislabeled ~2,628 low-quality completions as score 10.