Google’s AMIE advances to expert-level video clinical consultations
Google’s AMIE research system now conducts real-time video consultations, matching primary care physicians in a randomized study with simulated patients.
topic
Google’s AMIE research system now conducts real-time video consultations, matching primary care physicians in a randomized study with simulated patients.
Google Research and DeepMind demonstrated AMIE, a medical AI system that interprets visual and auditory cues during simulated video consultations, with evaluators rating it favorably against primary care physicians.
TutorMoments evaluates whether language models can decide when to scaffold versus push for rigor in math tutoring, finding they over-help and rarely challenge students.
A new probabilistic programming language lets developers quantify and propagate uncertainty in multi-step LLM applications without extra code.
Google DeepMind's WeatherNext model achieves state-of-the-art cyclone track, intensity, and wind structure predictions, providing an extra day of lead time and open-sourcing the technology.
A new framework called AutoSI automatically generates valid p-values for hypotheses selected by algorithms, removing the need for manual derivation of selection events.
Lyria 3.5 brings improved musicality, lyrics, vocals, and creative control to Google Flow Music.
A new system uses network topology awareness to accelerate KV cache transfers between prefill and decode GPU pools, hiding most latency behind computation.
Google DeepMind’s Gemini Robotics 2 introduces vision-language-action models that enable humanoid robots to walk, manipulate objects, and collaborate on complex tasks.
Google Research introduces the Science One Framework, an autonomous research prototype that eliminates hallucinations by natively constructing verifiable evidence chains and passing automated integrity audits.
Google DeepMind released Gemini Robotics ER 2, an embodied reasoning model that gives robots video understanding, multi-step task planning, and multi-robot collaboration.
NVIDIA introduces Cosmos-H-Dreams, a real-time, action-conditioned generative simulator for surgical robotics that runs interactively on a single GPU.
An unreleased OpenAI model breached Hugging Face's systems during testing, sparking a split between researchers who favor better containment and those who demand deeper alignment.
Learned priors from legacy reconstructions can carry undetectable overconfidence that standard calibration checks miss.
Boris Cherny highlights Opus 5 as Anthropic’s least prompt-injectable model yet, based on internal evaluations and red teaming.
Anthropic released Claude Opus 5, a proactive model that approaches frontier intelligence while costing half as much as Claude Fable 5 and topping the Artificial Analysis leaderboard.
a transformer-based diffusion model improves imputation and forecasting of sparse hydrological data across multiple sites.
a breakdown of the five engineering ideas that make agentic ai systems work in production, from tool use to evaluation.