The Polygenic Index (PGI) Repository

We compute polygenic indices for a wide range of traits and make them available as variables you download from the data providers.

47 phenotypes, single-trait and multi-trait (MTAG) PGIs
11 datasets that may be useful to researchers
Updated regularly, with additional PGIs and datasets

How to get and use the data

The PGIs are distributed by the data providers, not by us. Most researchers need the first four steps only.

Step 1

Check coverage

Confirm that your phenotype and your dataset are both in the Repository before anything else.

Step 2

Request the dataset

Each participating dataset has its own application procedure and data-use agreement. The PGIs arrive as variables you download from the provider alongside the rest of the data.

Step 3

Correct for measurement error

Every PGI is measured with error. Left uncorrected, that error attenuates your estimates.

Step 4

Cite the Repository

Cite Becker et al. (2021), together with the GWAS that went into the single-trait or multi-trait input for each PGI you use.

Optional

If you need the summary statistics or PGI weights

The datasets already carry the PGIs themselves, so most users never need this step. If you want the underlying GWAS and MTAG summary statistics or the LDpred weights, they are on our download portal.

What is available

Citing the Repository

Include this citation in any publication based on the Repository PGIs or the measurement-error-corrected estimator, along with the citations for the GWAS included in the single-trait or multi-trait input GWAS for the PGI.

Becker et al. 2021

Becker, J., Burik, C.A.P., Goldman, G., Wang, N., Jayashankar, H., Bennett, M., Belsky, D.W., Karlsson Linnér, R., Ahlskog, R., Kleinman, A., Hinds, D.A., 23andMe Research Group, Caspi, A., Corcoran, D.L., Moffitt, T.E., Poulton, R., Sugden, K., Williams, B.S., Harris, K.M., Steptoe, A., Ajnakina, O., Milani, L., Esko, T., Iacono, W.G., McGue, T., Magnusson, P.K.E., Mallard, T.T., Harden, K.P., Tucker-Drob, E.M., Herd, P., Freese, J., Young, A., Beauchamp, J.P., Koellinger, P.D., Oskarsson, S., Johannesson, M., Visscher, P.M., Meyer, M.N., Laibson, D., Cesarini, D., Benjamin, D.J., Turley, P., and Okbay, A. (2021). Resource Profile and User Guide of the Polygenic Index Repository. Nature Human Behaviour, 5, 1744–1758. doi:10.1038/s41562-021-01119-3

Read the paper: Becker et al. (2021) in Nature Human Behaviour.

Why a shared repository The four roadblocks this resource removes

Constructing PGIs is time-consuming

Building indices from individual genotype data takes real effort, even for researchers trained to work with large datasets.

The largest summary statistics are the hardest to obtain

Prediction accuracy increases with the sample size of the underlying GWAS, but privacy and IRB restrictions create administrative hurdles that force researchers to trade the benefit of a larger sample against the cost of clearing them.

Sample overlap causes overfitting

Publicly available summary statistics are sometimes based on a discovery sample that includes the target cohort, or close relatives of its members. That overlap can lead to highly misleading results.

Different methods are not comparable

Because researchers construct PGIs from summary statistics in different ways, results from different studies are hard to compare and interpret.

Summary statistics and PGI weights

For each phenotype in the Repository, we report GWAS and MTAG summary statistics and PGI (LDpred) weights for all SNPs from the largest discovery sample for that analysis, unless the sample includes 23andMe.

SNP-level summary statistics from analyses based entirely or in part on 23andMe data can only be reported for up to 10,000 SNPs. So if the largest GWAS or MTAG analysis for a phenotype includes 23andMe, we report summary statistics for only the genome-wide significant SNPs from that analysis, and additionally report summary statistics for all SNPs from the largest analysis excluding 23andMe.

To maximise prediction accuracy we meta-analysed summary statistics from multiple sources, including several novel large-scale GWASs conducted in UK Biobank and the personal genomics company 23andMe. Almost all PGIs in our initial release therefore perform at least as well as currently available PGIs in terms of prediction accuracy.

Find these data at thessgac.com

Measurement-error-corrected estimator

In Becker et al. (2021) we propose an approach that improves the interpretability and comparability of research based on PGIs. In place of ordinary least squares, we derive an estimator that corrects for errors-in-variables bias. It produces coefficients in units of the standardised additive SNP factor, which has a more meaningful interpretation than units of some particular PGI.

The Python command-line tool implementing the estimator

Interested in participating in the Repository?

We answer researcher queries about coverage and access directly.

contact@theSSGAC.org