Convex Markets / Datasets / Preference data

Dataset catalog

Filter by type, access, and pricing. Specs show before you open the product page.

10 results
S
External · PRF-3009

Skywork-Reward-Preference-80K

77,016 curated chosen/rejected chat preference pairs subsampled from public sources (HelpSteer2, OffsetBias, WildGuard, Magpie DPO series) and used to train the Skywork-Reward reward models.

Type Preference dataVolume 77016 preference pairs (chosen/rejected chat pairs)Format ParquetAccess PUBLIC LICENSE
Free
open dataset (HF)
View
H
External · PRF-3005

HelpSteer2

NVIDIA's open (CC-BY-4.0) helpfulness dataset of 21,362 human-annotated prompt-response samples rated on five attributes, with a dedicated human pairwise-preference split for training reward models and DPO.

Type Preference dataVolume 21,362 annotated prompt-response samples (20,324 train + 1,038 validation)Format JSON Lines (gzipped .jsonl.gz; loadable via HF datasets)Access PUBLIC LICENSE
Free
open dataset (HF), CC-BY-4.0
View
P
External · PRF-3004

PKU-SafeRLHF

A large human-annotated safety-preference dataset where each question has two model responses ranked separately for helpfulness and harmlessness, plus per-response safety meta-labels across 19 harm categories.

Type Preference dataVolume 83.4K preference entries (Q-A pairs, each with two ranked responses)Format JSONL (per-model train.jsonl / test.jsonl; auto-converted to Hugging Face Parquet)Access PUBLIC LICENSE
Free
open dataset (HF)
View
U
External · PRF-3001

UltraFeedback

A large-scale, fine-grained preference dataset of ~64k prompts, each with 4 model completions rated by GPT-4 across four aspects, for training reward and critique models.

Type Preference dataVolume 63967 prompts (train split rows; 256k completions total)Format Parquet (HF datasets); JSON in source repoAccess PUBLIC LICENSE
Free
open dataset (HF)
View
IO
External · PRF-3008

Intel Orca DPO Pairs

A ~12.9K-example preference dataset in Direct Preference Optimization (DPO) format, derived from Open-Orca/OpenOrca, pairing a 'chosen' and 'rejected' response for each instruction prompt.

Type Preference dataVolume 12,859 preference pairs (DPO triples: prompt + chosen + rejected)Format JSONL (single file orca_rlhf.jsonl; also auto-converted to Parquet on Hugging Face)Access PUBLIC LICENSE
Free
open-source license (Apache-2.0)
View
N(
External · SFT-2103

Nectar (7-wise Ranked Preference Data)

182,954 chat prompts each with 7 ranked responses (GPT-4-judged) — the preference dataset behind Berkeley's Starling reward model.

Type Preference dataVolume 182954 promptsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
CA
External · PRF-3006

Chatbot Arena Conversations

33K real pairwise human preference votes from the LMSYS Chatbot Arena, where users compared two anonymous LLMs side by side and chose the better response.

Type Preference dataVolume 33000 pairwise conversations (human preference votes)Format ParquetAccess GATED
Free
open dataset (HF, gated - must accept terms)
View
SH
External · PRF-3003

Stanford Human Preferences (SHP)

385,563 collective human preference pairs mined from 18 Reddit Q&A subreddits, where the more-helpful of two comments is inferred from Reddit scores and timestamps.

Type Preference dataVolume 385,563 preference pairsFormat JSON (one JSONL/JSON file per subreddit split; Parquet auto-conversion on HF)Access PUBLIC LICENSE
Free
open dataset (HF)
View
WC
External · PRF-3007

WebGPT Comparisons

19,578 human-scored answer-pair comparisons from OpenAI's WebGPT project, used to train the WebGPT long-form question-answering reward model.

Type Preference dataVolume 19578 preference comparisons (answer pairs)Format Parquet (Hugging Face auto-converted) / JSONL (original OpenAI source)Access PUBLIC LICENSE
Free
open dataset (HF)
View
AH
External · PRF-3002

Anthropic HH-RLHF

Anthropic's human preference dataset of chosen/rejected assistant dialogue pairs for training helpful and harmless (HH) reward models via RLHF.

Type Preference dataVolume 169,352 preference comparison pairs (chosen/rejected), default config: 160,800 train + 8,552 testFormat JSON Lines (.jsonl.gz) in the GitHub/HF repo; Parquet (auto-converted) via the Hugging Face datasets-serverAccess PUBLIC LICENSE
Free
open dataset (HF) / MIT license
View