Dataset catalog
Filter by type, access, and pricing. Specs show before you open the product page.
PettingZoo (Farama)
The standard Python API and reference environment suite for multi-agent RL, covering Atari (multiplayer), Classic board/card games, MPE, SISL, and Butterfly.
Procgen Benchmark
OpenAI's Procgen Benchmark is a suite of 16 procedurally-generated, Atari-like Gym/Gym3 reinforcement-learning environments built to measure sample efficiency and generalization in RL.
BabyAI
Grid-world instruction-following platform with a compositional synthetic 'Baby Language'; 19 levels of increasing difficulty, now part of Minigrid.
Jericho
Open-source Python RL environment from Microsoft Research that connects agents to 57 human-made interactive-fiction text games via a modified Frotz Z-machine interpreter, with reward defined as in-game score deltas.
Overcooked-AI
A two-agent gridworld cooking environment for benchmarking human-AI and multi-agent coordination, where agents cooperatively prepare and deliver soups for a shared sparse reward.
DROP
A crowdsourced reading-comprehension benchmark whose questions require discrete reasoning (addition, counting, sorting, comparison) over Wikipedia-derived paragraphs.
HotpotQA
A Wikipedia-based multi-hop question-answering dataset of 113,000+ QA pairs that require reasoning across multiple supporting documents and provide sentence-level supporting facts.
TextWorld
Microsoft's sandbox engine for procedurally generating and playing text-adventure games to train and evaluate RL agents; underlies ALFWorld.
OpenR1-Math-220k
220k competition math problems, each with 2-4 DeepSeek-R1 chain-of-thought traces verified by Math-Verify and Llama-3.3-70B; text reasoning, no code.
NuminaMath-1.5
~896k competition-math problems (Chinese high-school to IMO level) with chain-of-thought solutions; improved successor to NuminaMath-CoT.