Convex Markets / Datasets / SFT dataset / Stanford Alpaca (and Alpaca-Cleaned)
SKU SFT-2101 · Sold by External

Stanford Alpaca (and Alpaca-Cleaned)

Product specifications

SKUSFT-2101
Data typeSFT dataset
Volume52002 examples
Size on disk24.3 MB (Parquet)
FormatParquet (Hugging Face)
Access modelPUBLIC LICENSE
PricingFree · open-source license
Quality score
LicenseCC BY-NC 4.0
Stanford Alpaca is a 52,002-example instruction-following dataset produced with a modified Self-Instruct pipeline using OpenAI's text-davinci-003. Each row is a single-turn instruction/input/output triple plus a pre-formatted prompt 'text' field. The widely used yahma/alpaca-cleaned mirror (51,760 rows) corrects hallucinations, empty outputs, and merged-instruction errors. Note a license discrepancy: the original tatsu-lab/alpaca is tagged CC BY-NC 4.0, while the alpaca-cleaned mirror is tagged CC BY-4.0 on Hugging Face despite deriving from OpenAI outputs.