Convex Markets / Datasets / All departments

Dataset catalog

Filter by type, access, and pricing. Specs show before you open the product page.

110 results
S
External · SFT-2005

SlimOrca

Open-Orca's ~518k-row curated subset of OpenOrca — GPT-4 FLAN reasoning traces with GPT-4 verification against human annotations.

Type SFT datasetVolume 517982 rowsFormat JSONLAccess PUBLIC LICENSE
Free
open-source license
View
N(
External · SFT-2103

Nectar (7-wise Ranked Preference Data)

182,954 chat prompts each with 7 ranked responses (GPT-4-judged) — the preference dataset behind Berkeley's Starling reward model.

Type Preference dataVolume 182954 promptsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
A(
External · AGT-4001

AgentInstruct (AgentTuning)

1,866 reward-filtered, ReAct-style multi-turn agent interaction trajectories (GPT-4 generated) spanning six real-world agent tasks, released to fine-tune the AgentLM models from the AgentTuning paper.

Type Agent tracesVolume 1866 trajectories (across 6 task splits: os 195, db 538, alfworld 336, webshop 351, kg 324, mind2web 122)Format Parquet (ShareGPT-style conversation records: conversations list + id)Access PUBLIC LICENSE
Free
open dataset (HF)
View
L
External · DEM-6006

Language-Table

A large suite of human-collected, language-conditioned tabletop block-manipulation demonstrations from Robotics at Google, where a robot pushes colored blocks in response to natural-language instructions, released in RLDS/TFDS format.

Type DemonstrationsVolume 442,226 real-robot demonstration episodes (language_table split; 1,639,544 episodes total across all 9 released variants)Format RLDS / TFDS (TFRecord); episodes as sequences of stepsAccess PUBLIC LICENSE
Free
open-source license (Apache-2.0)
View
T(
External · AGT-4002

ToolBench (ToolLLM)

126,486 instruction-tuning instances pairing real user queries over 16,464 real-world RapidAPI APIs with ChatGPT-generated DFSDT (depth-first search decision tree) tool-use solution paths and final answers.

Type Agent tracesVolume 126486 instruction-solution-path instances (over 3451 tools / 16464 RapidAPI APIs)Format JSONAccess PUBLIC LICENSE
Free
open-source license (Apache-2.0)
View
U2
External · SFT-2004

UltraChat 200k

HuggingFaceH4's heavily filtered subset of UltraChat (~208k train_sft dialogues) used to train the Zephyr chat models.

Type SFT datasetVolume 207865 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
GF
External · AGT-4006

Glaive Function Calling v2

112,960 synthetic multi-turn chat conversations pairing a SYSTEM tool-schema prompt with assistant function-call turns, released by Glaive AI on Hugging Face under Apache-2.0.

Type Agent tracesVolume 112960 chat samples (multi-turn function-calling conversations)Format JSON (single glaive-function-calling-v2.json file; two string fields: system, chat)Access PUBLIC LICENSE
Free
open dataset (HF)
View
BV
External · DEM-6002

BridgeData V2

60,096 real WidowX 250 robot manipulation trajectories (teleoperated + scripted) across 24 environments and 13 skills, each labeled with natural-language task instructions.

Type DemonstrationsVolume 60096 trajectoriesFormat Per-trajectory raw files (obs_dict.pkl, policy_out.pkl, agent_data.pkl, lang.txt, images0/im_*.jpg 640x480 JPEG); also NumPy and TFRecord/RLDS (TFDS, 256x256) conversionsAccess PUBLIC LICENSE
Free
open dataset (CC-BY-4.0)
View
M
External · SFT-2203

MetaMathQA

395k math QA pairs bootstrapped from GSM8K and MATH via rephrasing, self-verification and backward (FOBAR) augmentation; text solutions only.

Type SFT datasetVolume 395k QA pairsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
M
External · SFT-2207

MathInstruct

262k math instruction examples combining chain-of-thought and program-of-thought (executable Python) rationales compiled from 13 datasets.

Type SFT datasetVolume 262k instruction-response examplesFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
Ev
External · SFT-2302

Evol-CodeAlpaca v1

111K English code instruction/output pairs — an Evol-Instruct augmentation of CodeAlpaca-20k using GPT-4 across 10 evolution strategies.

Type SFT datasetVolume 111272 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
CA
External · PRF-3006

Chatbot Arena Conversations

33K real pairwise human preference votes from the LMSYS Chatbot Arena, where users compared two anonymous LLMs side by side and chose the better response.

Type Preference dataVolume 33000 pairwise conversations (human preference votes)Format ParquetAccess GATED
Free
open dataset (HF, gated - must accept terms)
View
D
External · SFT-2108

Dolphin

Open Orca-style FLAN reproduction: ~892K FLANv2 completions from GPT-4 plus ~2.84M from GPT-3.5, filtered to remove refusals/alignment.

Type SFT datasetVolume 3731947 examplesFormat JSONL (Parquet auto-conversion available)Access PUBLIC LICENSE
Free
open-source license
View
WE
External · SFT-2102

WizardLM Evol-Instruct V2 196k

Complexity-evolved (Evol-Instruct) instruction conversations derived from Alpaca/ShareGPT for WizardLM SFT; HF split ships 143K rows.

Type SFT datasetVolume 143000 examplesFormat JSON (Parquet auto-conversion available)Access PUBLIC LICENSE
Free
open-source license
View
O
External · SFT-2405

OpenOrca

An open collection of ~2.94M FLAN-Collection instructions paired with GPT-4/GPT-3.5 'Orca-style' reasoning-trace responses for supervised instruction fine-tuning.

Type SFT datasetVolume 2,942,029 records (instruction-response pairs, single train split)Format ParquetAccess PUBLIC LICENSE
Free
open dataset (HF)
View
E
External · SFT-2303

Evol-Instruct-Code-80k-v1

78K code instruction/output pairs — an open reproduction of WizardCoder's Evol-Instruct-Code (CodeAlpaca run through 3 evolution rounds).

Type SFT datasetVolume 78264 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
T
External · AGT-4004

ToolAlpaca

A corpus of ~3,900 generalized tool-use agent trajectories over 400+ real-world APIs, generated by a multi-agent (user / assistant / tool-executor) simulation and formatted as ReAct-style Thought/Action/Action Input/Observation traces.

Type Agent tracesVolume 3,938 tool-use instances (from 400+ APIs across 50 categories; repo train_data.json contains 4,255 instances over 468 tool records)Format JSONAccess PUBLIC LICENSE
Free
open-source license (Apache-2.0)
View
S
External · SFT-2402

Self-Instruct

Machine-generated instruction-following dataset created by bootstrapping GPT-3 from 175 seed tasks, released as ~82K prompt/completion demonstrations for instruction tuning.

Type SFT datasetVolume 82612 instances (prompt–completion pairs) in the self_instruct configFormat Text; HF configs served as prompt/completion string pairs (Parquet via datasets-server); original repo ships JSON/JSONLAccess PUBLIC LICENSE
Free
open dataset (HF)
View
L
External · SFT-2008

LIMA

GAIR's 1,030-example gated dataset of high-quality curated prompts and responses used in the 'Less Is More for Alignment' study.

Type SFT datasetVolume 1030 rowsFormat Parquet (Hugging Face)Access GATED
Free
open-source license
View
A/
External · DEM-6007

ALOHA / ACT Demonstrations

Human-teleoperated bimanual fine-manipulation demonstrations collected with the low-cost open-source ALOHA hardware for the ACT paper, recorded as multi-camera RGB video plus 14-DoF joint states/actions in per-episode HDF5 files.

Type DemonstrationsVolume 50 human-teleoperated demonstrations per task (Thread Velcro: 100; 4 simulated task configs of 50 episodes each also released)Format HDF5 (one .hdf5 file per episode)Access PUBLIC LICENSE
Free
open-source license
View