
[Explainer] Where Does Your Y Chromosome Come From? Reading East Asian Paternal Migrations Through O, N, C and F
2026年4月1日
China’s Donghulin “Unknown Ancient Human Lineage” Study Draws International Archaeology Coverage
2026年4月6日The 1000 Chinese Pangenome: Westlake University Builds the Largest Human Pangenome to Date, Revealing 13% of Previously Unknown Sequences
On 1 April 2026, Nature published a milestone study: a team led by Professor Jian Yang of Westlake University, together with Professor Xian Shen of Wenzhou Medical University, built the largest and most comprehensive human pangenome resource to date—the 1000 Chinese Pangenome (1KCP)—systematically revealing the genomic diversity of the Chinese population.
What is a pangenome?
If each person’s genome is a book of life written in about 3 billion base pairs, the traditional human reference genome is like its standard edition. That standard edition, however, was built mainly from European genomes and cannot fully represent the genetic diversity of all populations. A pangenome is not a single version but the collection of genome sequences from a population, capturing the full spectrum of genetic variation and better reflecting the true nature of the human genome.
Human genomics research has long depended on European-based references. Even the Human Pangenome Reference Consortium (HPRC) included only 3 Chinese samples among its 46 global individuals—far from representing the diversity of 1.4 billion Chinese people. As a result, East Asian–specific rare variants are often missed or misclassified, directly affecting diagnostic accuracy and drug development.
Unprecedented scale and an innovative method
The team developed a pangenome-guided integrated assembly (PIGA) approach and constructed 1,116 diploid genome assemblies of Chinese individuals—55 de novo assemblies and 1,061 pangenome-guided assemblies—from a health-examination cohort in Wenzhou, greatly exceeding the sample size of previous pangenome efforts.
A genomic new world: 400 million novel sequences
Compared with the international reference genomes GRCh38 and CHM13, the 1KCP pangenome uncovered 405.3 million base pairs of novel sequences—13% of the human genome—never recorded before. More strikingly, 26.2 million bp of these have clear functional roles, spanning gene-coding regions and regulatory elements, revealing a wealth of new genetic components that may affect the health of Chinese people.
A comprehensive family portrait of genetic variation
The study systematically identified 35.4 million small variants, more than 110,000 structural variants (SVs), 485,000 tandem repeats, and 860,000 nested variants (secondary variants hidden inside complex SVs). Notably, 33.3% of the SVs were discovered for the first time, and 83.5% are rare variants (population frequency ≤1%)—often associated with inherited disease yet easily missed in small studies.
Direct evidence for medical applications
The team dissected variant–disease relationships: 5,239 SVs directly affect exons of protein-coding genes (an average of 450 per Chinese individual, 9 of them rare); 2,427 tandem-repeat expansion events were identified, linked to neurodegenerative diseases such as Huntington’s disease and spinocerebellar ataxia; and 735 gene clusters with SVs were found, significantly enriched in blood and immune functions, reflecting human adaptation to diverse pathogens.
Significance
The 1KCP project fills the genetic reference gap for precision medicine in the Chinese population, providing a population-specific framework for disease risk prediction, pharmacogenomics, and personalized treatment, with an open 1KCP data portal for the research community. Two decades after China contributed 1% of the sequencing in the Human Genome Project, Chinese scientists have now independently completed this landmark study—a leap from follower to leader in the field.
Paper: The 1000 Chinese Pangenome empowers medical and population genetics (Nature, 2026).




