A new approach to quantifying the personal information contained in gut metagenome data

[ad_1]

In a recent study published in Natural microbiologythe researchers used shotgun sequencing to extract human deoxyribonucleic acid (DNA) readings in fecal samples from 343 Japanese individuals comprising the main dataset for this study.

They used this gut metagenome data to reconstruct personal information. Some study participants also provided whole genome sequencing (WGS) data for ultra-deep sequencing analysis of the metagenome.

Study: Reconstruction of personal information from human genome reads in gut metagenome sequencing data.  Image Credit: KaterynaKon/Shutterstock.comStudy: Reconstruction of personal information from human genome reads in gut metagenome sequencing data. Image Credit: KaterynaKon/Shutterstock.com

Background

Knowledge about the human microbiome, the microorganisms inhabiting the human body, has grown significantly over the past decade, thanks to rapid advances in technologies such as shotgun metagenome sequencing.

This technology allows the sequencing of the non-bacterial component of microbiome samples, including host DNA. For example, in fecal samples, the amount of host DNA is less than 10% but is removed to protect donor privacy.

The human germline genotype in metagenome data is important to allow re-identification of individuals. However, researchers and donors should recognize that it is highly confidential, so sharing it with the community requires careful consideration.

In addition to the ethical concerns of sharing this data, it is necessary to understand that if human reads in the metagenome data are not removed prior to submission, what type of personal information (e.g. gender and ancestry ) could this data help to recover?

Additionally, human reads in gut metagenome data could be a good resource for stool-based forensics, robust variant calling, and disease risk estimates based on polygenic risk scores (e.g., the Type 2 diabetes).

Since these data could help to quantitatively and accurately reconstruct genotype information, they could complement human WGS data.

About the study

In the current study, the researchers applied a few human reads into gut metagenome data from the study’s main dataset to reconstruct personal information, including genetic sex and ancestry. To predict the genetic sex and ancestry of these 343 individuals, they used depth of sex chromosome sequencing and the modified likelihood score-based method, respectively.

Additionally, researchers have developed methods to re-identify a person from a genotype dataset. In addition, they combined two harmonized genotype calling approaches, direct calling of rare variants and two-step imputation of common variants, to reconstruct genotypes.

The main dataset for the study included 343 Japanese participants, while the validation dataset for the genetic sex prediction analysis included 113 Japanese individuals.

The multi-ancestry dataset, which helped the researchers validate the ancestry prediction analysis, included 73 individuals of different nationalities, including samples from individuals in New Delhi, India.

The female and male participants in each dataset were 196 and 147, 65 and 48, and 25 and 48, respectively. Similarly, the age range for these three datasets was 20-88, 20-81 and 20 to 61, respectively.

Results and conclusion

Since the human reads in the gut metagenome data were consistently derived from all chromosomes, the read depth of the X chromosome was almost double in females and that of the Y chromosome in males.

So, in a logistic regression analysis, when researchers applied a Y:X chromosome read depth ratio of 0.43 to the validation dataset, which correctly predicted the genetic sex of 97.3 % of study samples.

In human microbiome and genetic research, the feasibility of sex prediction using human gut metagenome data could help weed out mislabeled samples.

The study’s analysis also helped the researchers remarkably predict the ancestry of 98.3% of individuals using data from the 1000 Genomes Project (1KG) as a benchmark.

However, the likelihood score-based method often misclassified South Asian (SAS) samples into American (AMR) and European (EUR), especially when the number of human reads was low. This is understandable because the genetic diversity of the SAS population is complex.

The likelihood score-based method also effectively used data from low-coverage genomic areas demonstrating the quantitative power of gut metagenome data to re-identify individuals and succeeded in re-identifying 93.3% of individuals.

Despite ethical concerns, the re-identification method used in this study could aid in the quality control of multi-omics datasets including gut metagenome and human germline genotype data.

Furthermore, the authors succeeded in reconstructing common genome-wide variants using genomic approaches. Historically, researchers have used stool samples as a source of germline genomes for wild and domestic animals, but not for humans.

Thus, further development of appropriate methodologies could help to effectively use the human genome in gut metagenome data and benefit animal research.

Nevertheless, the study remarkably demonstrated that optimized methods could help reconstruct personal information from human reads in gut metagenome data.

Additionally, the results of this study could serve as a guiding resource for designing best practices for utilizing the data already accumulated on the human gut metagenome.

Sources

1/ https://Google.com/

2/ https://www.news-medical.net/news/20230517/A-novel-approach-to-quantify-personal-information-contained-within-gut-metagenome-data.aspx

The mention sources can contact us to remove/changing this article

[ad_2]

Leave a Reply

Your email address will not be published. Required fields are marked *

Related Posts