source: Google Research: Transfer learning for genomic prediction in underrepresented populations
level: research
google research studied how transfer learning from european biobank data affects polygenic risk scores in japanese populations. they used uk biobank and biobank japan across eight traits. the goal was to find when adding european samples helps or hurts prediction. they tested elastic net, meta-analysis, and prs-csx models. the key finding is that more out-of-population data is not always better. for small target samples, european data boosts accuracy. but once the target cohort grows, that benefit disappears and can reverse.
the crossover point is around 15,000 japanese samples. below that, pooling european data improves prediction. above it, training only on target data works better. the exact point depends on trait genetic correlation. for traits like bmi, which are genetically similar across populations, european data helps up to 25,000 to 40,000 samples. for lipids and blood glucose, which have population-specific genetics, the benefit fades much earlier. prs-csx needed over 25,000 samples to match simpler elastic net models.
the study used eight clinical traits including bmi, blood pressure, blood cell counts, cholesterol, and glucose. heritability in uk biobank ranged from 0.07 to 0.28. meta-analysis helped most for population-specific traits at small sample sizes. prs-csx performed worse than elastic net below 25,000 samples but matched or exceeded it near 100,000. the results suggest that building diverse local biobanks is essential. model choice should depend on trait heritability and available sample size.
why it matters: for ai in genomics, this shows that blindly adding diverse training data can reduce model accuracy, so practitioners must tune data sources to target population size and trait genetics.
source: Google Research: Transfer learning for genomic prediction in underrepresented populations