Convex Markets / Datasets / RL environments / SWE-bench / SWE-bench Verified
SKU RLE-1203 · Sold by External

SWE-bench / SWE-bench Verified

Product specifications

SKURLE-1203
Data typeRL environments
Volume2294 task instances
Size on diskNot published by source
FormatHugging Face dataset (Parquet/JSON) + Docker
Access modelPUBLIC LICENSE
PricingFree · open-source license
Quality score
LicenseMIT
SWE-bench evaluates LLMs/agents on real-world software issues collected from GitHub: given a repository and an issue, the agent must produce a patch that passes hidden regression tests. The full test set has 2,294 instances; key subsets include SWE-bench Verified (500 human-validated), Lite (300), Multimodal, and Multilingual, plus a large train split without executable environments. Evaluation runs in sandboxed Docker with FAIL_TO_PASS/PASS_TO_PASS test checks. Data is distributed as public Hugging Face datasets (SWE-bench, _Verified, _Lite, _Multimodal).