Dataset catalog
Filter by type, access, and pricing. Specs show before you open the product page.
xLAM Function-Calling 60k
60,000 verifiable single/multi/parallel function-calling instances (query + available tools + reference tool-call answers) generated and triple-verified by Salesforce's APIGen pipeline.
Gorilla APIBench
Instruction-to-API-call dataset over HuggingFace, TorchHub, and TensorHub APIs, used to train and evaluate LLMs that write correct API calls (the Gorilla / APIBench benchmark).
ToolACE
11,300 synthetic multi-turn function-calling dialogues generated over a self-evolved pool of 26,507 APIs and filtered by a dual-layer (rule-based + model-based) verification pipeline.
WebLINX
Expert demonstrations of conversational, multi-turn website navigation, where a navigator agent must predict the next web action (click/say/load/submit/change) from dialogue history and the page DOM.
Mind2Web
2,350 human-demonstrated web-navigation task trajectories across 137 real websites, each pairing a natural-language instruction with a step-by-step sequence of DOM-grounded CLICK/TYPE/SELECT actions and full HTML snapshots.
AgentInstruct (AgentTuning)
1,866 reward-filtered, ReAct-style multi-turn agent interaction trajectories (GPT-4 generated) spanning six real-world agent tasks, released to fine-tune the AgentLM models from the AgentTuning paper.
ToolBench (ToolLLM)
126,486 instruction-tuning instances pairing real user queries over 16,464 real-world RapidAPI APIs with ChatGPT-generated DFSDT (depth-first search decision tree) tool-use solution paths and final answers.
Glaive Function Calling v2
112,960 synthetic multi-turn chat conversations pairing a SYSTEM tool-schema prompt with assistant function-call turns, released by Glaive AI on Hugging Face under Apache-2.0.
ToolAlpaca
A corpus of ~3,900 generalized tool-use agent trajectories over 400+ real-world APIs, generated by a multi-agent (user / assistant / tool-executor) simulation and formatted as ReAct-style Thought/Action/Action Input/Observation traces.