SKU SFT-2405 · Sold by External

OpenOrca

Product specifications

SKUSFT-2405
Data typeSFT dataset
Volume2,942,029 records (instruction-response pairs, single train split)
Size on disk~2.87 GB (Parquet files; ~5.05 GB uncompressed in-memory per HF size API)
FormatParquet
Access modelPUBLIC LICENSE
PricingFree · open dataset (HF)
Quality score
LicenseMIT
OpenOrca is an open dataset of augmented FLAN Collection data whose responses were generated by querying GPT-4 and GPT-3.5, aligned as closely as possible to the data distributions described in Microsoft's Orca paper (arXiv:2306.02707). Each record contains a source FLAN instruction, an Orca-style system prompt, and a model-generated response that often includes detailed step-by-step reasoning ('reasoning traces'). The HF card describes the collection as ~1M GPT-4 completions plus ~3.2M GPT-3.5 completions, distributed as two Parquet files (1M-GPT4-Augmented.parquet and 3_5M-GPT3_5-Augmented.parquet); the currently hosted single train split reports 2,942,029 rows via the Hugging Face size/rows API. The data is a single unsplit train set spanning four FLAN submixes identified in the id prefix (niv, t0, cot, flan). It was released by the Open-Orca / AlignmentLab.ai team and used to train models such as Mistral-7B-OpenOrca and OpenOrca-Platypus2-13B. The card tags the dataset MIT, but responses were produced with OpenAI models, so downstream commercial use warrants review.