Updated
Updated · Google Research · Sep 3
Google Study Finds PRS Transfer Learning Fades Beyond 15,000 Japanese Samples
Updated
Updated · Google Research · Sep 3

Google Study Finds PRS Transfer Learning Fades Beyond 15,000 Japanese Samples

1 articles · Updated · Google Research · Sep 3

Summary

  • Google Research found European-trained polygenic risk models help underrepresented populations only when local datasets are small, with gains in Biobank Japan reversing once target training samples reach about 15,000.
  • Eight traits tested across UK Biobank and Biobank Japan showed the crossover depends on genetic similarity: conserved traits such as BMI benefited from pooled European data until roughly 25,000-40,000 samples, while HDL, LDL and blood glucose lost accuracy earlier.
  • More than 5,000 UK Biobank samples could already reduce HDL prediction performance after that 15,000-sample crossover, suggesting out-of-population data can become harmful as target-population evidence strengthens.
  • Meta-analysis of European and Japanese GWAS improved population-specific traits at smaller Japanese sample sizes, while PRS-CSx generally lagged elastic-net models below 25,000 samples and only matched or beat them near 100,000.
  • The study argues that better genetic risk prediction for underrepresented groups will require larger local biobanks and model choices tailored to trait architecture and sample size, not simply adding more European data.

Insights

Could using vast European genetic data actually harm risk predictions for other populations once local datasets reach a certain size?
Why might genetic risk scores with poor statistical accuracy still save more lives in historically underrepresented populations?
If genetic risk scores fail to predict lifestyle responses, are we fundamentally misunderstanding their true clinical utility?