Convex Markets / Datasets / All departments

Dataset catalog

Filter by type, access, and pricing. Specs show before you open the product page.

60 results
T3
External · SFT-2001

Tulu 3 SFT Mixture

Allen AI's 939k-row supervised fine-tuning mixture combining public and synthetic instruction data used to train the Tulu 3 models.

Type SFT datasetVolume 939344 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
OA
External · SFT-2105

Orca AgentInstruct 1M v1

1,046,410 synthetic instruction/response conversations generated by Microsoft's AgentInstruct agentic pipeline across 15 task-type splits.

Type SFT datasetVolume 1046410 examplesFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
G(
External · RLE-1406

Gymnasium (Farama)

The standard Python RL API and reference environment suite (successor to OpenAI Gym) covering Classic Control, Box2D, Toy Text, MuJoCo, and Atari.

Type RL environmentsVolume dozens reference environmentsFormat Python (pip: gymnasium; Gym reset/step API)Access PUBLIC LICENSE
Free
open-source license
View
O
External · SFT-2204

OpenMathInstruct-2

14M math problem-solution pairs (GSM8K/MATH augmentation) generated by Llama-3.1-405B-Instruct; text chain-of-thought solutions.

Type SFT datasetVolume 14M problem-solution pairsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
S/
External · RLE-1203

SWE-bench / SWE-bench Verified

Benchmark of real GitHub-issue tasks (2,294 full; 500 human-verified) evaluated by FAIL_TO_PASS/PASS_TO_PASS unit tests in Docker.

Type RL environmentsVolume 2294 task instancesFormat Hugging Face dataset (Parquet/JSON) + DockerAccess PUBLIC LICENSE
Free
open-source license
View
A
External · RLE-1205

AppWorld

Simulated world of 9 apps and 457 APIs with 750 interactive coding tasks evaluated by state-based unit tests.

Type RL environmentsVolume 750 tasksFormat Python package + SQLite DBs + JSON task specsAccess PUBLIC LICENSE
Free
open-source license
View
V
External · RLE-1102

VisualWebArena

Multimodal, visually grounded web-agent benchmark: 910 tasks over self-hosted Classifieds, Shopping and Reddit sites.

Type RL environmentsVolume 910 tasks (Classifieds 234, Shopping 466, Reddit 210)Format JSON task configs + Python (gym) + self-hosted Docker sitesAccess PUBLIC LICENSE
Free
open-source license
View
W/
External · RLE-1104

WorkArena / WorkArena++ (ServiceNow)

Enterprise knowledge-work benchmark on live ServiceNow instances: 33 L1 atomic tasks (19,912 instances) plus 682 L2/L3 compositional tasks.

Type RL environmentsVolume 33 L1 atomic tasks (19,912 instances) + 682 WorkArena++ L2/L3 tasksFormat Python package (pip install browsergym-workarena) + live ServiceNow instance + PlaywrightAccess GATED
Free
open-source license
View
I
External · SFT-2104

Infinity-Instruct

BAAI's large-scale open SFT collection; multi-config (7M chat, 3M foundational, plus dated Gen sets). Gated on Hugging Face.

Type SFT datasetVolume 7449106 examplesFormat Parquet (Hugging Face)Access GATED
Free
open-source license
View
t(
External · RLE-1206

tau-bench (τ-bench)

Tool-agent-user benchmark of 165 customer-service tasks (retail 115, airline 50) with policy-following and pass@k evaluation.

Type RL environmentsVolume 165 tasks (retail 115 + airline 50)Format Python package + JSON domain dataAccess PUBLIC LICENSE
Free
open-source license
View
W
External · SFT-2107

WildChat-1M

Real user–ChatGPT (GPT-3.5/GPT-4) conversations collected by Ai2 with metadata and moderation labels; current train split ~838K conversations.

Type SFT datasetVolume 837989 conversationsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
W
External · RLE-1101

WebArena

Realistic self-hosted web environment: 812 long-horizon tasks over shopping, forum, GitLab, CMS and maps with execution-based evaluation.

Type RL environmentsVolume 812 tasks (from 241 intent templates)Format JSON task configs + Python (gym) + self-hosted Docker sitesAccess PUBLIC LICENSE
Free
open-source license
View
C
External · SFT-2306

CodeFeedback-Filtered-Instruction

156.5K high-quality single-turn code instructions filtered (complexity 4-5 via Qwen-72B-Chat) from four open code-instruction sources; M-A-P.

Type SFT datasetVolume 156526 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
o
External · SFT-2206

orca-math-word-problems-200k

200k grade-school math word problems with GPT-4-Turbo-generated worked solutions; English, text explanations only.

Type SFT datasetVolume 200k word-problem QA pairsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
MO
External · SFT-2301

Magicoder OSS-Instruct 75K & Evol-Instruct 110K

ISE-UIUC's paired Magicoder code-instruction datasets: OSS-Instruct (75K, seeded from open-source snippets) and Evol-Instruct (110K, decontaminated evol-codealpaca).

Type SFT datasetVolume 186380 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
O2
External · SFT-2002

OpenHermes 2.5

Teknium's ~1M-row compilation of primarily GPT-4-generated instruction, chat, coding and reasoning data in ShareGPT (from/value) format.

Type SFT datasetVolume 1001551 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
N(
External · SFT-2103

Nectar (7-wise Ranked Preference Data)

182,954 chat prompts each with 7 ranked responses (GPT-4-judged) — the preference dataset behind Berkeley's Starling reward model.

Type Preference dataVolume 182954 promptsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
S
External · SFT-2005

SlimOrca

Open-Orca's ~518k-row curated subset of OpenOrca — GPT-4 FLAN reasoning traces with GPT-4 verification against human annotations.

Type SFT datasetVolume 517982 rowsFormat JSONLAccess PUBLIC LICENSE
Free
open-source license
View
U2
External · SFT-2004

UltraChat 200k

HuggingFaceH4's heavily filtered subset of UltraChat (~208k train_sft dialogues) used to train the Zephyr chat models.

Type SFT datasetVolume 207865 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
M
External · SFT-2203

MetaMathQA

395k math QA pairs bootstrapped from GSM8K and MATH via rephrasing, self-verification and backward (FOBAR) augmentation; text solutions only.

Type SFT datasetVolume 395k QA pairsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View