SKU RLE-1204 · Sold by External
Terminal-Bench
Product specifications
| SKU | RLE-1204 |
|---|---|
| Data type | RL environments |
| Volume | 89 tasks |
| Size on disk | Not published by source |
| Format | Docker + YAML/dir task specs + Python harness |
| Access model | PUBLIC LICENSE |
| Pricing | Free · open-source license |
| Quality score | — |
| License | Apache-2.0 |
Terminal-Bench evaluates AI agents on hard, realistic command-line tasks (software engineering, security, scientific computing, data science, ML, sysadmin, debugging) executed inside Docker containers. Agents issue tmux keystrokes / bash commands via a ReAct-style loop (Terminus reference agent) and are scored by outcome-based test scripts checking the final container state. Terminal-Bench 2.0 comprises 89 curated tasks (from 229 created by 93 contributors); 2.1 is the current release. Each task is a directory with an instruction, a Dockerized environment, a test script, and a reference solution.