SKU PRF-3006 · Sold by External
Chatbot Arena Conversations
Product specifications
| SKU | PRF-3006 |
|---|---|
| Data type | Preference data |
| Volume | 33000 pairwise conversations (human preference votes) |
| Size on disk | 41.6 MB (Parquet download); ~81.2 MB uncompressed |
| Format | Parquet |
| Access model | GATED |
| Pricing | Free · open dataset (HF, gated - must accept terms) |
| Quality score | — |
| License | CC-BY-4.0 (user prompts) / CC-BY-NC-4.0 (model outputs) - dual/mixed |
Chatbot Arena Conversations is a preference dataset of 33,000 cleaned, real pairwise human votes collected from the LMSYS Chatbot Arena between April and June 2023, contributed by over 13,000 unique users (IP addresses) across roughly 20 large language models. In each record a user chatted with two anonymous models side by side and voted which response was better, producing a single pairwise human preference label stored in the 'winner' field (observed values model_a, model_b, tie; 'tie (bothbad)' is also documented). Each row contains the two model names, both full conversations in OpenAI message format, an anonymized judge ID, the turn count, detected language, a Unix timestamp, and safety metadata (OpenAI moderation API output plus RoBERTa-large and T5-large toxicity tags). The corpus is distributed as a single Parquet file on Hugging Face behind a gated agreement that requires accepting disclaimer terms about potentially unsafe content. It was released by LMSYS Org alongside the paper 'Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena' (arXiv:2306.05685, NeurIPS 2023 Datasets and Benchmarks).