The Polygenic Index (PGI) Repository
We compute polygenic indices for a wide range of traits and make them available as variables you download from the data providers.
How to get the data
1 Check coverage Phenotype and dataset lists 2 Request the dataset The PGIs come as variables from the data provider 3 Correct for measurement error pgi_correct 4 Cite the Repository Becker et al. 2021 — Summary statistics and PGI weights Only if you need the weights themselves. Most users can skip this.How to get and use the data
The PGIs are distributed by the data providers, not by us. Most researchers need the first four steps only.
Step 1
Check coverage
Confirm that your phenotype and your dataset are both in the Repository before anything else.
Step 2
Request the dataset
Each participating dataset has its own application procedure and data-use agreement. The PGIs arrive as variables you download from the provider alongside the rest of the data.
Step 3
Correct for measurement error
Every PGI is measured with error. Left uncorrected, that error attenuates your estimates.
Step 4
Cite the Repository
Cite Becker et al. (2021), together with the GWAS that went into the single-trait or multi-trait input for each PGI you use.
Optional
If you need the summary statistics or PGI weights
The datasets already carry the PGIs themselves, so most users never need this step. If you want the underlying GWAS and MTAG summary statistics or the LDpred weights, they are on our download portal.
Resources
Plain-language FAQs explain how PGIs should — and should not — be interpreted and used.
Citing the Repository
Include this citation in any publication based on the Repository PGIs or the measurement-error-corrected estimator, along with the citations for the GWAS included in the single-trait or multi-trait input GWAS for the PGI.
Becker, J., Burik, C.A.P., Goldman, G., Wang, N., Jayashankar, H., Bennett, M., Belsky, D.W., Karlsson Linnér, R., Ahlskog, R., Kleinman, A., Hinds, D.A., 23andMe Research Group, Caspi, A., Corcoran, D.L., Moffitt, T.E., Poulton, R., Sugden, K., Williams, B.S., Harris, K.M., Steptoe, A., Ajnakina, O., Milani, L., Esko, T., Iacono, W.G., McGue, T., Magnusson, P.K.E., Mallard, T.T., Harden, K.P., Tucker-Drob, E.M., Herd, P., Freese, J., Young, A., Beauchamp, J.P., Koellinger, P.D., Oskarsson, S., Johannesson, M., Visscher, P.M., Meyer, M.N., Laibson, D., Cesarini, D., Benjamin, D.J., Turley, P., and Okbay, A. (2021). Resource Profile and User Guide of the Polygenic Index Repository. Nature Human Behaviour, 5, 1744–1758. doi:10.1038/s41562-021-01119-3
Read the paper: Becker et al. (2021) in Nature Human Behaviour.
Why a shared repository The four roadblocks this resource removes
Constructing PGIs is time-consuming
Building indices from individual genotype data takes real effort, even for researchers trained to work with large datasets.
The largest summary statistics are the hardest to obtain
Prediction accuracy increases with the sample size of the underlying GWAS, but privacy and IRB restrictions create administrative hurdles that force researchers to trade the benefit of a larger sample against the cost of clearing them.
Sample overlap causes overfitting
Publicly available summary statistics are sometimes based on a discovery sample that includes the target cohort, or close relatives of its members. That overlap can lead to highly misleading results.
Different methods are not comparable
Because researchers construct PGIs from summary statistics in different ways, results from different studies are hard to compare and interpret.
Summary statistics and PGI weights
For each phenotype in the Repository, we report GWAS and MTAG summary statistics and PGI (LDpred) weights for all SNPs from the largest discovery sample for that analysis, unless the sample includes 23andMe.
SNP-level summary statistics from analyses based entirely or in part on 23andMe data can only be reported for up to 10,000 SNPs. So if the largest GWAS or MTAG analysis for a phenotype includes 23andMe, we report summary statistics for only the genome-wide significant SNPs from that analysis, and additionally report summary statistics for all SNPs from the largest analysis excluding 23andMe.
To maximise prediction accuracy we meta-analysed summary statistics from multiple sources, including several novel large-scale GWASs conducted in UK Biobank and the personal genomics company 23andMe. Almost all PGIs in our initial release therefore perform at least as well as currently available PGIs in terms of prediction accuracy.
Measurement-error-corrected estimator
In Becker et al. (2021) we propose an approach that improves the interpretability and comparability of research based on PGIs. In place of ordinary least squares, we derive an estimator that corrects for errors-in-variables bias. It produces coefficients in units of the standardised additive SNP factor, which has a more meaningful interpretation than units of some particular PGI.
Interested in participating in the Repository?
We answer researcher queries about coverage and access directly.