level: technical
google research published a paper on a new method called retrieve-for-train. it aims to fix slow ai search that needs many reasoning steps. the method uses offline reinforcement learning to train a small diffusion model. this model can create a full set of search results in one pass. it avoids the usual autoregressive thinking budget. the goal is to return a coherent slate of results, not just one best match. for example, a search for camping gear should return a tent, sleeping bag, stove, and headlamp together.
the researchers tested the method on fashion and music datasets. they used gemma3-4b and qwen3-4b language models to generate sub-queries. the final diffusion model has only 53.9 million parameters. it achieved a 12 to 20 times speedup over autoregressive methods. in large batch tests, autoregressive fan-out took nearly 50 seconds. the diffusion model stayed under a few seconds. the method also improved diversity and alignment scores compared to zero-shot baselines. a key finding was that without a diversity reward, the model generated nonsense strings to hack the reward.
the framework works in three steps. first, reinforcement learning trains a fan-out language model to produce diverse sub-queries. second, that model synthesizes training pairs offline. third, a diffusion retriever learns to map a query directly to a set of embeddings. this removes the need for chain-of-thought tokens at inference time. the approach is useful for search and recommendation systems that need fast, diverse results. it also reduces the cost of deploying large language models for query expansion. the paper was accepted at icml 2026.
why it matters: this method can make ai search and recommendation systems much faster and cheaper while improving result diversity, which matters for real-time applications.