Showing posts with label slow clock. Show all posts
Showing posts with label slow clock. Show all posts

Monday, October 27, 2014

Why the surprising pattern of no genetic continuity between people living in the same area but from different periods of time? Think the flu virus!

I used three slides as shown below to illustrate the idea of informative DNAs in my talk in last month’s workshop on genome and evolution in Naples, Italy.

The antigenic sites in human influenza A virus mutate and turn over quickly, which is critical for their survival or escape from human neutralizing antibodies and hence responsible for flu epidemics. As shown in Figure 1, two amino acid positions in hemagglutinin (156 and 145, panel a and b) turned over several times within a 30 year period, while two others (138 and 194, panel c and d) stayed largely unchanged (Figure from Shih et al, 2007). 

The flu results illustrate two important points with regard to evolutionary dynamics of a genome that have so far been grossly overlooked by the evolution and popgen field. First, fast evolving or less conserved DNAs are also functional rather than neutral as they are essential for quick adaptive needs in response to fast changing environments. Second, fast evolving DNAs turn over quickly and can be shown to violate the infinite sites model.  Hence, they cannot be used for phylogenetic inference. If one uses the fast changing sites in a flu virus to infer the phylogenetic relationship of the virus isolates responsible for different epidemics in a past period of say 10 years, one would reach the absurd conclusion that each epidemic was caused by a distinct type of flu virus with no genetic continuity among them rather than just minor variations of the same type.

Mutation rates in humans are of course much slower than that in a flu virus. But just like a flu virus, there are also fast and slow changing sites (Figure 2). The time scales are different but the principle is the same.  The fast changing sites may turn over every few thousand years and in fact make up the majority of the observed variant sites in humans when properly examined by us (Figure 3). This is why the field of ancient DNA kept producing the absurd pattern of no genetic continuity between people living in the same area but from different periods of time. All of the published analyses have simply used the wrong sites that are equivalent to the fast changing antigenic sites in a flu virus. What one should be using are sites with very slow mutation rates, like 1 mutation every 50,000 years. We have been busy reinterpreting the published DNAs for several years now and hope to submit our work soon.


Figure 1. (a and b) Frequency changes at residue sites 156 (a) and 145 (b) were highly dynamic. (c and d) Sites 138 (c) and 194 (d) did not undergo major frequency change over time.




Figure 2. A priori model of evolutionary dynamics of human genomic DNAs.




Figure 3. Difference between slow and fast evolving sites. Shown are a piece of homologous DNA in three different individuals or species. In the fast evolving DNAs making up the vast majority of human genome, there is obvious and verifiable violation of the infinite sites model. These DNAs have abundant overlapped mutant sites where independent mutations have occurred on the same site in different individuals or species. 


Ref.


Shih, C-C., Hsiao, T-C., Ho, M-S., and Li, W-H. (2007) Simultaneous amino acid substitutions at antigenic sites drive influenza A hemagglutinin evolution. Proc Natl Acad Sci U S A. 104:6283-6288.

Sunday, November 11, 2012

Evidence of natural selection of Y chr invalidates dating of human divergence history using Y chr markers

I attended the American Society of Human Genetics 2012 meeting last week.  The following work presented by a Berkeley lab at the meeting shows that the whole Y chr is under strong purifying natural selection and shows extremely low diversity.  Thus the markers on Y chr are not neutral and should not be used for phylogeny inferences.  If two different human populations share some hypolotypes in Y chr, it does not indicate shared ancestry but rather some sort of common selection.  We have been saying in recent years that nearly all human genome variations are not neutral and are under natural selection.  There are essentially no junks.  To date human evolution history, one must use slow evolving neutral sequences as we have been advocating.  All existing literature on human history using Y chr or mtDNA or any other sequences are mistaken. 

There are only one scientific theory known so far that advocates no junk DNAs, i.e., the MGD hypothesis.  Much work in recent years have essentially killed the junk DNA concept, most recently by the ENCODE finding of at least 80% human genome being functional.  But the theory that predicts the neutral and junk DNA concept still remains to be overthrown.  The MGD represents the best chance to explain nearly 100% functional genome and to supercede the neutral theory. 

Abundant selection explains low diversity on human Y chromosomes. M. Wilson Sayres1,2, K. Lohmueller1,2, R. Nielsen1,2 1) Integrative Biology, University of California, Berkeley, Berkeley, CA; 2) Statistics, University of California, Berkeley, Berkeley, CA.

The human Y chromosome exhibits levels of diversity that are significantly lower than expected under neutral population genetic theory. Variance in male reproductive success (reducing the effective population size of males relative to females) has recently been proposed as an alternative neutral model to explain reduced diversity on the Y relative to mtDNA. Generally Y chromosomes are not included in whole genome analyses, so explicit tests of this hypothesis have yet to be conducted. Here we show that neutral models with unequal male and female effective population sizes are not consistent with observed genome-wide diversity on autosomes, X, Y and mtDNA across completely sequenced males. Instead, a model including selection is needed to explain the departure of observed Y diversity from expectations. We found that models with similar estimates of the strength of background selection can explain diversity for both the Y chromosome and mitochondrial genomes. Our results suggest that strong selection is necessary for explaining the evolutionary history of the human Y chromosome, and argue against the concept of the "junk" Y chromosome .

 

Thursday, December 30, 2010

Others have also noticed the difference between fast and slow evolving genes in phylogeny inference

The anthropology blogger Dienekes at http://dienekes.blogspot.com had a recent post on Dec 30, 2010 noticing the major difference between slow and fast evolving genes in calculating the time of origin for humans.

It is reassuring to see that others have come to the same conclusion as I have regarding the difference between slow and fast genes. The mere existence of the difference between slow and fast evolving genes is not predicted by the modern evolution theory, which is at best incomplete since it does not cover the situation when maximum genetic distance has been reached. But maximum genetic diversity hypothesis covers it and is therefore a more complete theory. The slow clock method based on this hypothesis has produced phylogeny results dramatically different from the populars ones. The results show that humans separated from the pongid clade 17.3 million years ago. For details, read a preprint here, http://precedings.nature.com/documents/3794/version/1

Conclusion: all existing molecular interpretations of life trees are based on fast evolving genes and are therefore incorrect. We will not have a correct interpretation until we have reanalyzed the data using our slow clock method based on the MGD hypothesis.

Thursday, July 15, 2010

New fossil finds on ape-monkey common ancestor supports my molecular dating while contradicts all others'

New Oligocene primate from Saudi Arabia and the divergence of apes and Old World monkeys. Iyad S. Zalmout, William J. Sanders, Laura M. MacLatchy, Gregg F. Gunnell, Yahya A. Al-Mufarreh, Mohammad A. Ali, Abdul-Azziz H. Nasser, Abdu M. Al-Masari, Salih A. Al-Sobhi, Ayman O. Nadhra, Adel H. Matari, Jeffrey A. Wilson & Philip D. Gingerich. Nature 466, 360–364 (15 July 2010)

This new paper on a common ancestor of ape-monkey from 29-28 million years ago in Nature this week fully supports my molecular dating as found in this preprint here (http://precedings.nature.com/documents/3794/version/1), while contradicts all other previous dating results on the ape-monkey divergence time. The dating in my paper is 29.7 million years (MY). Dating results by others using completely different, in my view flawed, methodology are self-conflicting and either too late (23 MY) or too early (30-35 MY). Any methodology that can turn solid factual data like DNA into conflicting interpretation of reality has of course self-proven itself false.

My paper has been through a number of submissions and I have not seen a valid criticism on the key points of the paper. The paper is now in the hands of a journal editor and below is part of my cover letter explaining my paper, which should help people understand the difference between my method/result and others’ and why others’ are incorrect from both a theoretical point of view and a practical point of view that they are contradictory to the new fossil finds.

Molecular phylogeny methods are based on the neutral theory. But the neutral theory should never have been invented in the first place for macroevolution if people had not overlooked the overlap feature of the genetic equidistance result that originally inspired the molecular clock and in turn the neutral theory. More on this exceptional new finding, see my newly published paper, Huang S (2010) The overlap feature of the genetic equidistance result, a fundamental biological phenomenon overlooked for nearly half of a century. Biological Theory 5: 40-52. http://www.mitpressjournals.org/doi/abs/10.1162/BIOT_a_00021

Since the neutral theory has no concept of a maximum distance, all these methods include a large amount of sequence alignment information that are non-informative and hence contribute to the high noise level that can sometimes overwhelm the signal. Also, these methods require false or uncertain assumptions that treat macroevolution the same as microevolution. It is thus fully expected that these methods cannot possibly produce a true molecular phylogeny of macroevolution. In fact, they have all self-proven themselves incorrect by repeatedly turning solid factual data into conflicting interpretations of reality, one of which must be false. The data in molecular phylogeny is just sequence facts and cannot possibly be wrong so long one is not making sequencing errors. Thus the only way to produce a false result or conflicting results in molecular phylogeny is through an incorrect method, including any method that does not have any correct means and principles of identifying only the informative data. One good example of endless conflicting results produced by the existing methods is the position of tarsiers, which some studies group with prosimians whereas others with simians, despite the fact that all these studies used the same kind of method but just different set of genes. A correct method should have ways of selecting the informative data and either produce only correct results or no results if informative data are not available.

The manuscript here used a new method ‘the slow clock’ to resolve key questions of primate phylogeny. Since the method has no uncertain assumptions and uses only informative genes, the slow evolving ones not yet reaching maximum genetic distance, it is immune to turning factual data into false or conflicting interpretation of reality. Indeed, the new primate molecular phylogeny here produced by the new method matches closely with the original interpretation of the fossil records by paleontologists.

You don’t really need to be reminded of this of course but I still suggest that you must use the highest standard of science to compare my story versus the existing theory/methodology. One must judge a story only by how contradiction-free it is. It is really the minimum scientific standard and the only practical way to distinguish a scientific truth from a religious belief. Thus, you only need to give me one single contradiction to my theory/methodology for me to withdraw my manuscript. That by the way should also be the only scientific way to reject the manuscript. I am eager to have the reviewers to help me either improve the manuscript or thrash it.

The following novel results in my manuscript are the contradictions to the existing theory/methodology but are not to mine. You and the reviewers must express your views on these results or offer an explanation if it happens to be different from mine. 1) Chimpanzee is closer to orangutan or gorilla than human is in DNA. 2) Slow and fast evolving genes produce different phylogeny. 3) The clock/neutral theory is a mistaken interpretation of the equidistance result of Margoliash. 4) Given 3), the existing interpretation of ape-human relationship is based on false theory and cannot possibly be true. One simply cannot imagine that the field has been all along on the right track when a major mistake had been committed right from the start. Neither can one imagine that the field can continue business as usual now that the mistake has finally been caught after nearly half of a century. Essentially no conclusion or interpretation on macroevolution from the molecular evolution field in the past half of a century can be regarded as correct or conclusive. 5) The fact that octopus is closer to human than cockle is, or that bird is closer to human than snake is in DNA contradicts the existing theory/methodology. The fact that chimpanzee is closer to human than orangutan or gorilla is therefore cannot be used to infer closer genealogy to human, just like one cannot infer closer genealogy between bird and human than between snake and human.

Finally, the newest version of my paper has a lay abstract as follows:

Primate phylogeny: molecular evidence for a pongid clade excluding humans and a prosimian clade containing tarsiers

Author Summary:

Molecular phylogeny methods are based on the neutral theory of evolution, which was originally inspired in a large part by the genetic equidistance result of Margoliash in 1963. But the neutral theory should never have been invented in the first place for macroevolution if people had not overlooked the overlap feature of the genetic equidistance result. Therefore, the field of molecular phylogeny has been on the wrong track all along, and a complete reevaluation of all molecular phylogeny results is in order. The maximum genetic diversity hypothesis is a more coherent and complete account of evolution, and was here used to resolve key questions of primate phylogeny. The analysis shows that humans are genetically more distant to orangutans than African apes are and separated from the pongid clade 17.3 million years ago. Also, tarsiers are genetically closer to lorises than simian primates are, suggesting a tarsier-loris clade to the exclusion of simian primates. The validity and internal coherence of the primate phylogeny here were independently verified. The results as a whole show a remarkable and unprecedented concordance between molecules and fossils.

Friday, November 13, 2009

Real reason for the endless production of conflicting results on tarsiers

A couple of weeks ago, a new paper by Chatterjee et al. appeared on primate phylogeny that groups tarsiers with prosimians.

“Estimating the phylogeny and divergence times of primates using a supermatrix approach” Helen J Chatterjee, Simon YW Ho, Ian Barnes, and Colin Groves
BMC Evolutionary Biology 2009, 9:259 doi:10.1186/1471-2148-9-259

I sent the following comment titled “Real reason for the endless production of conflicting results on tarsiers” to the Journal’s website:

On the position of tarsiers, Chatterjee et al wrote: “The majority of molecular evidence supports the latter grouping [4,10-13] (grouping tarsiers with higher primates), although a large number of molecular studies still provide support for the Prosimii concept [14-18].”

When a method or technique can lead to two opposite results repeatedly and seemingly endlessly while only one of the two can be true, it is time to ask whether something is fundamentally missing with our method (all existing popular methods are slightly different from one another but are fundamentally the same kind). Let us start from the very beginning and examine the assumptions for our method. The key assumption for all sequence similarity based methods is that sequence dissimilarity always correlates with time of divergence. Well, is this true? We don’t have to be a specialist to know that this is sometimes true and sometimes not. Thus, for our method to be able to produce accurate and uncertainty-free result, it must take into account the reality that sequence dissimilarity sometimes does not correlate with time of divergence. Many of the sequence comparisons are not informative and should and must be excluded from our method. When they are not as is the case with all existing methods, they contribute to the high noise level that can sometimes overwhelm the signal. It is by accident that these methods sometimes give correct results and sometimes wrong results and no one knows why the difference or when to view such a result correct and final. Therefore, we have a peculiar non-scientific situation: no one is taking anyone else’s results as the final say. Never mind that we only have one true phylogeny of life on Earth. Once you know it, it is done and no more work needed. The existing methods are perfect for keeping some of us employed forever but will never give us truth. Truth is not judged by a quantitative difference in the number of studies that support it versus those against it. The correct method should produce zero number of studies that is against truth or should be immune to the production of conflicting results.

Data + method = result. The data here in molecular phylogeny is just sequence facts and cannot possibly be wrong so long one is not making sequencing errors. Thus the only way to produce a false result or conflicting results in phylogeny is through an incorrect method. Since all existing methods are perfectly capable of producing false results and have all in practice produced false results or conflicting results, it is another simple proof that the existing popular methods are simply incorrect.

By the own admission of the leading experts, the existing popular methods are flawed in the sense that they can easily produce incorrect results that are totally out of the hands of the scientists:
“Unlike the case in physics, the predictive power of a model in biology is quite low. It seems to us that if the prediction (e.g., a phylogenetic tree reconstructed) of a model is correct in 80% of the cases, it is a good model at least at the present time.” From Masatoshi Nei and Sudhir Kumar, 2000, Molecular Evolution and Phylogenetics, (p85):

When a result is only 80% certain, it can be completely wrong. We either know or we don’t know. Knowing with 80% certainty or anything less than 100% means we don’t know. We are much better off without it because it often leads the non-specialists into the wrong idea that we know with 100% certainty. Does not everyone in academic think that we are 100% certain that chimp is closest to humans when in fact we are only 80% certain and can therefore be completely wrong? When they then act and work based on that knowledge (they have been doing just that for years now), should we feel perfectly comfortable for misleading them into that?

Of course, nothing we know says that biology has to be different from physics. The present situation merely means that we have much to learn. When we know better, we should be able to have a model or method that is correct in 100% of the cases. Until then, some of what we are doing is just kidding ourselves. I have now offered the slow clock method as the best candidate for a method that takes into account all reality and is capable of 100% certainty (1). While the result of Chatterjee et al., like many others, does support my result on tarsiers using the slow clock method (1), I do not view their result as confirming mine, because their method is flawed. By using the same kind of method, another group could easily produce a result opposite of theirs by just picking a new set of genes (this of course has been done many times already). I of course do not view such result as valid contradiction to mine just like I do not view the result of Chatterjee et al. as valid support.

A flawed method automatically qualifies its result as meaningless, regardless whether the result happens to be consistent with reality or not. The definition of a flawed method is simply that which can turn a perfectly solid set of factual data into a false interpretation of reality. Any method that has produced conflicting interpretations has of course automatically self-proven itself false. The present situation we have with tarsiers is just one of many that says flatly it is time for a fundamental change in our method of interpreting sequence data.

Ref:
Huang, S. (2009) Primate phylogeny: molecular evidence for a pongid clade excluding humans and a prosimian clade containing tarsiers. Available from Nature Precedings, http://hdl.handle.net/10101/npre.2009.3794.1