SKU SFT-2101 · Sold by External
Stanford Alpaca (and Alpaca-Cleaned)
Product specifications
| SKU | SFT-2101 |
|---|---|
| Data type | SFT dataset |
| Volume | 52002 examples |
| Size on disk | 24.3 MB (Parquet) |
| Format | Parquet (Hugging Face) |
| Access model | PUBLIC LICENSE |
| Pricing | Free · open-source license |
| Quality score | — |
| License | CC BY-NC 4.0 |
Stanford Alpaca is a 52,002-example instruction-following dataset produced with a modified Self-Instruct pipeline using OpenAI's text-davinci-003. Each row is a single-turn instruction/input/output triple plus a pre-formatted prompt 'text' field. The widely used yahma/alpaca-cleaned mirror (51,760 rows) corrects hallucinations, empty outputs, and merged-instruction errors. Note a license discrepancy: the original tatsu-lab/alpaca is tagged CC BY-NC 4.0, while the alpaca-cleaned mirror is tagged CC BY-4.0 on Hugging Face despite deriving from OpenAI outputs.