SKU PRF-3009 · Sold by External
Skywork-Reward-Preference-80K
Product specifications
| SKU | PRF-3009 |
|---|---|
| Data type | Preference data |
| Volume | 77016 preference pairs (chosen/rejected chat pairs) |
| Size on disk | ~209 MB (Parquet download) / ~416 MB uncompressed |
| Format | Parquet |
| Access model | PUBLIC LICENSE |
| Pricing | Free · open dataset (HF) |
| Quality score | — |
| License | No license published by source |
Skywork-Reward-Preference-80K-v0.2 is a subset of ~80K human/model preference pairs assembled by Skywork (Kunlun Inc.) to train the Skywork-Reward-Gemma-2-27B-v0.2 and Skywork-Reward-Llama-3.1-8B-v0.2 reward models. Each record is a (chosen, rejected) pair of multi-turn chat message lists plus a source label indicating which public dataset it came from. The pairs are subsampled (with no other modification) from HelpSteer2, OffsetBias, WildGuard (adversarial), and the Magpie DPO series (Ultra, Pro Llama-3.1, Pro, Air); Magpie samples were selected by average ArmoRM score and the WildGuard subset was additionally filtered by a Skywork reward model so that the chosen response scores higher than the rejected one. Version v0.2 is the decontaminated release: 4,957 magpie-ultra-v0.1 pairs with significant n-gram overlap against RewardBench evaluation prompts were removed. The dataset is a single Parquet train split of 77,016 rows totaling ~416 MB uncompressed (~209 MB download).