source: Google Research: ToolGrad: Efficient tool-use dataset generation with textual "gradients"
level: research
google research presented toolgrad at acl 2026, a new framework for generating tool-use datasets. unlike prior methods that start with a user query and search for a valid tool chain, toolgrad first generates a successful tool-use chain and then writes the matching user prompt. this answer-first approach uses textual gradients, feedback from an llm critic, to iteratively build complex api workflows. the framework has four modules: api proposer, api executors, api selector, and llm updater.
toolgrad was tested using the toolbench api database with over 16,000 real-world apis. compared to the query-first depth-first search baseline, toolgrad achieved a 99.8 percent pass rate in data generation, while requiring fewer optimization steps and lower cost. fine-tuning gemma-3 models on a small dataset of 500 toolgrad samples improved tool-use performance across 1b, 4b, and 12b parameter sizes. the 12b model scored 83.1 on the berkeley function calling leaderboard, nearly matching gemini-2.5-pro at 83.2 and outperforming gpt-5 at 74.4.
the toolgrad-500 dataset was generated using gemini-2.5-flash-lite, yet the fine-tuned gemma-3-12b model outperformed its teacher model, showing self-evolving capability. toolgrad-12b also beat open-source tool-use models like toolace and hammer-2.1-7b. the approach addresses cost and scalability bottlenecks in creating ground-truth tool-use data. future work may extend toolgrad to dynamic api ecosystems and continuous learning for personalization. this method could make training capable ai agents more economical for enterprise and everyday tasks.
why it matters: toolgrad reduces the cost and effort of creating high-quality tool-use training data, enabling smaller open models to match proprietary systems on tool-calling benchmarks.
source: Google Research: ToolGrad: Efficient tool-use dataset generation with textual "gradients"