source: google research: symptomai: towards a conversational ai agent for everyday symptom assessment

level: research

google research conducted a randomized national study with 13,917 participants to evaluate conversational ai agents for symptom assessment. participants described their symptoms to one of five gemini flash 2.0-based agents, which asked follow-up questions and generated a differential diagnosis. two weeks later, participants reported any diagnoses received from healthcare providers. a panel of three board-certified clinicians then reviewed the conversation transcripts and ranked the ai-generated diagnoses against those from other clinicians.

the clinicians preferred the ai's differential diagnoses over those from other clinicians in more than 50% of cases. the ai's top-5 accuracy was also higher, meaning its list of possible diagnoses more often included the actual diagnosis later confirmed by a healthcare provider. the study found that agents actively asking follow-up questions significantly outperformed a baseline where the ai did not prompt for more information. the ai's advantage was greatest in cases where clinicians were least confident in their own assessments.

researchers also correlated ai diagnoses with biosignal data from participants' fitbit devices. for those diagnosed with respiratory infections, physiological metrics like heart rate, respiration, and skin temperature showed clear shifts in the days before symptom reporting. this suggests ai symptom checkers could enable large-scale analysis of wearable data to identify disease patterns. the study is exploratory and all ai-generated labels were for research only, not clinical diagnoses.

why it matters: accurate ai symptom assessment could improve healthcare access and enable population-scale health research using wearable data.


source: google research: symptomai: towards a conversational ai agent for everyday symptom assessment