
Who’s Afraid of Genetic Ancestry?
2024年7月23日
Origins of the Korean Population: Insights from Mitochondrial DNA and Y-Chromosome Markers
2024年8月29日The Y Chromosome and Genetic Genealogy
Published: March 12, 2013. Author: Li Hui (Ministry of Education Key Laboratory of Contemporary Anthropology, Fudan University)
\n\n\n\nGenealogies are an important tool for remembering our ancestors and preserving blood ties; the records in one genealogy belong to members of a family sharing the same surname. Since surnames are mostly inherited from the father, and the Y chromosome is a genomic segment transmitted strictly from father to son, males sharing a surname tend to carry identical or closely related Y-chromosome types. Over the long course of history, genealogical documents are often lost for various reasons, and the Y chromosome—a genealogy written in the genome—can help us reconstruct the family trees that existed in reality. The stable point mutations (SNPs) on the Y chromosome are passed down forever through paternal descendants, allowing a reliable paternal genetic phylogeny to be built, while the faster-mutating microsatellite loci (STRs) can be used to estimate time. The Y chromosome can therefore be used to study the history of many surname clans, and even to resolve historical mysteries a thousand years old. Rebuilding the relationship among surnames, genealogies and the Y chromosome will surely become an important part of historical anthropology, and genetic genealogy will surely become a powerful tool for a harmonious society.
\n\n\n
In modern society, almost everyone has a surname. A surname is not merely a symbol: it carries rich cultural, historical and clan significance. Surnames, traced along lines of blood, record the origins of each family and are often used in research on ancestral roots. When people of the same surname meet, they often say, “We were one family 500 years ago,” and compiling a genealogy that sorts out the kinship among people of the same surname is the wish of many. For thousands of years, most surnames have been transmitted through the paternal line, and the Y chromosome of the human genome follows paternal inheritance even more strictly, so surnames and the Y chromosome correspond well in parallel. With the discovery of numerous genetic markers on the Y chromosome, using the Y chromosome to analyze relationships within same-surname populations—even among populations worldwide—will play an important role in molecular anthropology, and genetic genealogy will inevitably exert a major influence in modern society.
\n\n\n\nClan surnames and the paternal inheritance of the Y chromosome
\n\n\n\nPaternal kinship is the principal kinship recorded in genealogies. Although surnames generally follow paternal inheritance, they do not do so completely. In the Chinese social context, adoption, inheritance through other lines, uxorilocal marriage (matrilocal residence), and even outright name changes can weaken the association between a surname and paternal blood. Many circumstances that affect paternal kinship are not faithfully recorded in genealogies. On the other hand, most Chinese surnames originated in the various enfeoffed states of the Spring and Autumn period; when the common people of a state all took the state’s name as their surname, their bloodlines may well have been heterogeneous from the start, producing internal genetic heterogeneity within many large surnames. The same surname does not necessarily imply the same origin. Even so, when we focus not on a whole surname within a population but on clans with explicit historical records or even genealogies, the surname remains an excellent genetic marker.
\n\n\n\nUnlike surnames, the human Y chromosome directly represents paternal inheritance: it is always transmitted from father to son and is unaffected by any social, cultural or natural factors. The human body has 23 pairs of chromosomes. Of the 22 pairs of autosomes, each pair has one chromosome from the father and one from the mother; corresponding segments of the two chromosomes are exchanged during transmission, producing an admixture effect—what geneticists call recombination. The other pair, the sex chromosomes, consists of the X and Y chromosomes. In females, the X chromosomes are also paired, one from each parent, so they too cannot escape the effects of admixture. In males, however, there is only one X chromosome from the mother and one Y chromosome from the father—meaning that a male’s Y chromosome can come only from his father. The mode of inheritance of the sex chromosomes therefore dictates that the Y chromosome follows strict paternal inheritance (Figure 1).
\n\n\n\nCan recombination occur between the Y and X chromosomes? To answer this, one must first understand the structure of the Y chromosome. Human Y-chromosomal DNA contains roughly 60 million base pairs. About 5% at the two ends of the chromosome constitutes the pseudoautosomal regions, which recombine with the corresponding segments of the X chromosome during transmission; the main trunk, about 95%, is the non-recombining region, which recombines with no chromosome at all. This property of the Y-chromosome trunk guarantees that sons inherit the paternal Y-chromosome trunk intact, free from admixture, ensuring the strict paternal inheritance of the trunk—an unalterable genetic genealogy.
\n\n\n\nTherefore, when lost or unfaithfully recorded surname genealogies can no longer serve as reliable evidence for tracing ancestors, studying the types of the Y-chromosome trunk on the basis of modern molecular biological techniques is the best method for directly tracing paternal relationships among members of a clan surname—and the only means of verifying the paternal connection between ancestors and descendants and completing a genealogy. For example, by analyzing Y-chromosome features in the descendants of Cao Cao, we can learn about Cao Cao’s own Y-chromosome characteristics and about the degrees of kinship among modern descendants of the Cao family. In fact, over any period with reasonably reliable historical records, the association between a family’s surname and its paternal inheritance can be guaranteed, so a family surname and a fixed Y-chromosome type are transmitted together and closely linked.
\n\n\n
A Y chromosome that changes slowly and steadily
\n\n\n\nGeneration after generation of father-to-son transmission, the Y chromosome slowly accumulates changes. It is precisely the accumulation of mutations that makes Y chromosomes of more distantly related individuals differ more in the human paternal inheritance system. The individual differences created by mutations on the Y chromosome are mainly of two kinds: single-nucleotide polymorphisms (SNPs) and short tandem repeats (STRs). DNA molecules are made of four bases (A, T, C and G) linked in a certain order; an SNP is a change of base type at a single position. A given SNP on the Y chromosome generally has only two types in a population. An STR, by contrast, is a specific segment of the chromosome in which a unit of several bases repeats; different Y chromosomes often have different numbers of repeats at the same STR position. Because SNPs and STRs differ in mutation properties and mutation rates, they serve different purposes in analysis.
\n\n\n\nTo establish a paternal inheritance system, the most important precondition is that ancestral mutations can be stably preserved on the Y chromosomes of descendants. Because SNP mutations have an extremely low mutation rate, they can be preserved permanently in descendants: descendants can only accumulate new mutations on top of ancestral ones, and cannot lose the ancestral mutation signature. By comparing the Y-chromosome differences between humans and chimpanzees, and the degree of Y-chromosome difference within large families, the SNP mutation rate on the Y chromosome has been calculated: the probability of an SNP mutation at one chromosomal position for each male born is roughly one in 30 million. In fact, because of the conservation of the Y euchromatic region and the fact that, throughout human history, many men left no male descendants surviving to the present, the actual mutation rate in populations should be several orders of magnitude lower. We usually study the euchromatic region of roughly 30 million base pairs in the non-recombining part of the Y chromosome; at a mutation rate of one in 30 million per base pair, each male has on average one new mutation in this region. This new mutation appears randomly at any point in the Y euchromatic region. If a second mutation later occurs at the very same point, the mutation is lost from descendants and we can no longer determine the ancestral Y-chromosome mutation spectrum through descendants. But the probability that two mutations occur successively at the same point is, by the rules of probability, the square of one in 30 million—that is, one in 900 trillion—which, relative to the total human population since ancient times, is virtually zero. We can therefore say that in the vast majority of cases, SNP mutations that appeared on an ancestral Y chromosome can be found in descendants, and descendants can only add new mutations to the ancestral Y-chromosome mutation spectrum.
\n\n\n\nA combination of mutations formed by several SNPs is called a haplotype. In Figure 2, for example, five SNP mutations successively give rise to five haplotypes: type 1 is the ancestral type of the others, which are all descendant types. The ancestral type together with all descendant types is called a haplogroup. Theoretically, all Y chromosomes of a family belong to a single haplogroup, because all the males in it should descend from a single ancestor.
\n\n\n\nOf course, the concept of a haplogroup can be large or small. On a grand scale, all Y chromosomes in the world belong to one haplogroup, all descending from a single late archaic human male in East Africa more than 200,000 years ago. More finely, the world can be divided into 20 trunk haplogroups, numbered A through T (Figure 3). The oldest, haplogroups A and B, never left Africa; C and D first reached Australia and Asia; E reached Asia and then returned to Africa; F gave rise to haplogroups such as G, H, I and J that formed the European race in the West, and to haplogroup K, from which N, O, P and Q arose to form the Mongoloid race in the East—with O becoming the mainstream lineage of the Chinese and Q the mainstream lineage of Native Americans. The Y-chromosome phylogeny thus constructs a great family tree of all humanity.
\n\n\n\nA clock on the Y chromosome
\n\n\n\nUsing the stably inherited SNPs on the Y chromosome, we can construct unambiguous genetic relationships between individuals or families. Moreover, since SNPs have a stable mutation rate, when we count the number of mutational differences between the Y chromosomes of different people and divide that number by the rate, after conversion we can estimate the divergence time between the two Y chromosomes—this is the “molecular clock” that measures evolutionary time. However, because the SNP mutation rate is extremely low and the mutational differences between individuals are scattered across the Y chromosome, finding them requires whole-Y-chromosome sequencing, and whole-genome sequencing is still too costly for general application. This shortcoming is compensated by the other genetic marker on the Y chromosome, the STR.
\n\n\n\nSome STR loci are located at fixed positions on the Y chromosome. The repeat unit within each STR locus changes its copy number during transmission, and this change also proceeds at a fixed rate. The STR mutation rate is far higher than that of SNPs: in families, the probability of mutation at each STR locus per male born is about one in 300. In a typical Y-chromosome analysis, we survey 15 STR loci, with a combined mutation rate of about one in 20. The Y chromosome carries roughly 150 STRs with 4–6 base-pair repeats; if all STR loci were analyzed, the combined mutation rate would be about one in two. This high mutation rate is very favorable for estimating divergence times between different Y chromosomes, which is why STR loci have become the “clock” on the Y chromosome.
\n\n\n\nSTR mutation is bidirectional: copy numbers can increase or decrease. At the same STR locus, different individuals with the same ancestor may show different mutation directions and repeat numbers. As with SNPs, STRs at several different positions can also form haplotypes. Analyzing the diversity of STR haplotypes in a population allows us to calculate the time to the population’s common ancestor. Assuming that each STR mutation increases or decreases the repeat count by exactly one—the single-step mutation model—and that the population has a constant effective size, the approximate time of a particular Y-SNP can be derived from the formula t = −Ne × ln(1 − V/Ne × μ), where Ne is the effective population size, μ is the mutation rate, ln is the natural logarithm, V is the observed variance of a certain STR value in the population, and the calculated t is the number of generations elapsed; multiplying by the number of years per generation gives the time.
\n\n\n\nAt the combined Y-STR mutation rate of about one in two, almost every person can form a unique haplotype. However, because mutations occur step by step, individuals who are more closely related in the paternal line have more similar STR haplotypes; a surname transmitted purely through the paternal line should carry similar STR haplotypes. But because the STR mutation rate is unstable, and because of the effect of back mutations, the error in time estimates from STRs remains very large. For accurately analyzing the divergence time of Y-chromosome haplogroups, the mutation spectrum of whole-Y-chromosome SNPs is therefore still required; in this respect, the anthropology laboratory of Fudan University is at the forefront of the world. In theory, once a sufficient number of Y-chromosome SNPs and STRs is available, surveying the haplotypes of males within a surname clan would make it possible to clearly construct the Y-chromosome phylogenetic tree of the family—and even to compile a clear genetic genealogy.
\n\n\n\nPractical applications of Y-chromosome research on surname genealogies
\n\n\n\nMultiple studies confirm that surname transmission in various countries is relatively stable. Using the Y chromosome to examine historical cases of disputed family relationships has produced several successful examples. A notable one concerns Thomas Jefferson, the third president of the United States, who was accused of fathering children with a maidservant. By comparing the Y-chromosomal polymorphic sites of male descendants of Jefferson’s uncle and of the maidservant’s two sons, the conclusion was reached that Jefferson was the biological father of the maidservant’s youngest son. The Y chromosome can not only resolve centuries-old disputes but also reach back thousands of years and corroborate traditions recorded in the Bible. According to the Bible, the priests among the Jews descend by blood from Aaron, the first high priest of Judaism. Skorecki, himself a priest of Ashkenazi Jewish descent, noticed that he differed greatly in physical characteristics from a priest of Sephardic Jewish descent; together with Hammer, an expert on Y-chromosome research, he analyzed the haplotypes of Jewish priests using the Y-chromosomal polymorphic sites YAP and DYS19. The results showed that Ashkenazi and Sephardic Jewish priests are more closely related to each other than to non-priest Jews—that is, the priests trace back 3,300 years to a common paternal ancestor. The perfect correspondence between Y-chromosome analysis and the biblical account is genuinely astonishing.
\n\n\n
Many studies of the correlation between Chinese surnames and the Y chromosome have also been published. Analyses of Y-chromosomal genetic polymorphism among unrelated males with surnames such as Li, Wang and Zhang within the same region show that the Y-chromosomal genetic polymorphism of unrelated males sharing these surnames is rich, and does not differ significantly from the genetic diversity of unrelated Han males with different surnames. This shows that the large Han surnames have essentially no internal homogeneity, and that relevant Y-chromosome research can only be carried out within clearly defined surname clans. Clan genealogies can only be reconstructed through the Y chromosome, not inferred merely from shared surnames or shared ancestral residences.
\n\n\n\nThe internal heterogeneity of the large Han surnames has many possible causes. In an ideal scenario, each surname has a single origin—that is, the founder of the surname was one person or several people with the same Y-chromosome haplotype, and no interference (name changes, non-biological paternity, etc.) occurred during surname transmission; in that case a surname could be identified by a single SNP-STR haplotype. But most Chinese surnames did not have a single origin. In the Zhou dynasty, most surnames derived from enfeoffed states and later became family names. For example, the royal descendants of the state of Cao took the surname Cao, but the descendants of its servants could also take the surname Cao—indeed, all commoners within the whole state could. Since the origins of the people within the state of Cao were diverse, with all kinds of Y chromosomes, Chinese surnames on the whole are internally heterogeneous in paternal blood.
\n\n\n\nMoreover, just as Y-chromosome STR haplotypes evolve into more and more types over time, the longer a surname has gone through transmission, the more social interference it has suffered and the greater the differences it displays. In China, surnames have a history of nearly 5,000 years; their origins are complex, and they have been affected by name changes to avoid disasters, to observe taboos, through adoption, by imperial bestowal or demotion, and by minorities adopting Han surnames. As a simple example, 53 of China’s 100 largest surnames are said to derive from the surname Ji. Studying Chinese surnames is thus extremely difficult; but China also has a tradition of compiling genealogies, and Y-chromosome genetic genealogy research is of great help in clarifying these intricate blood relationships.
\n\n\n\nA genealogy is a special genre of book that records, in tabular form, the descent and reproduction of a family whose members share a common ancestor, centered on blood relationships and including other aspects of the family. In other words, those entered in a genealogy must share a common ancestor; even with the same surname, people of different ancestry cannot be compiled into the same genealogy. In China’s vast rural areas, people have long practiced the custom of same-surname settlement, and with the relatively small marriage radius, the same-surname population of a given region, as defined by genealogies, can be regarded as a paternally isolated population with identical or closely related Y chromosomes—an excellent research model for molecular anthropological analysis of Y-chromosomal DNA diversity.
\n\n\n\nSome genealogies, however, contain fabricated or borrowed content, so genealogical materials must be used with care. Yet before the irrefutable scientific evidence of Y-chromosome testing, any genealogy can be tested and corrected. The study of the association among surnames, genealogies and the Y chromosome will surely become a new powerful tool for the public to compile family trees, an important approach to studying the origins and evolution of the Chinese people, and a new chapter in historical anthropological research.
\n\n




