source: arXiv Artificial Intelligence: Iris: Climbing to the Search Frontier

level: research

researchers present iris-mini and iris-pro, two search agents trained at 35b-a3b and 397b-a17b scales. they also release the data pipeline and training recipe. tasks are built from web hyperlink structure. multi-hop chains are authored over an entity graph from a seed page and its out-links. every non-answer entity is rewritten into a descriptive reference, so no clue can be resolved by string matching. only questions that a reference model fails closed-book but solves with evidence are admitted.

these questions become trajectories, filtered at trajectory and turn level before supervised fine-tuning. the policy is then optimized by reinforcement learning against live search. the reward judge and observation summarizer run inside the training cluster. over-long rollouts are interrupted at the request level and resumed from their committed prefix at the next step. this design keeps training efficient while handling long search sessions.

the approach targets a known weakness: models that answer from memory alone. by forcing descriptive references and multi-hop reasoning, the agents must retrieve and connect evidence. the two scales suggest a trade-off between cost and capability. the release of the pipeline and recipe lets other teams reproduce or adapt the method. for ai and data science, this shows how to build search agents that generalize beyond memorized facts.

why it matters: this method creates search agents that must retrieve and reason over evidence, reducing reliance on memorized answers and improving factual reliability.


source: arXiv Artificial Intelligence: Iris: Climbing to the Search Frontier