SKU PRF-3001 · Sold by External
UltraFeedback
Product specifications
| SKU | PRF-3001 |
|---|---|
| Data type | Preference data |
| Volume | 63967 prompts (train split rows; 256k completions total) |
| Size on disk | 940 MB |
| Format | Parquet (HF datasets); JSON in source repo |
| Access model | PUBLIC LICENSE |
| Pricing | Free · open dataset (HF) |
| Quality score | — |
| License | MIT License |
UltraFeedback is a large-scale, fine-grained, diverse preference dataset built by OpenBMB (THUNLP) to train reward and critique models for RLHF. About 63,967 instructions were sampled from six public sources (UltraChat, ShareGPT, Evol-Instruct, TruthfulQA, FalseQA, and FLAN) and used to query a pool of 17 LLMs, sampling 4 responses per prompt for 256k completions in total. GPT-4 then annotates each completion under a fine-grained instruction covering four aspects (instruction-following, truthfulness, honesty, and helpfulness), producing both numerical ratings and textual rationales plus an overall_score. The card notes RLHF researchers can construct roughly 340k comparison pairs from these ratings. A December 2023 update corrected an overall_score bug that had mislabeled ~2,628 low-quality completions as score 10.