Wednesday, May 6, 2009

The last straw for Molecular clock

I sent the following email yesterday to a small groups of colleagues.

Dear colleagues with an interest in human evolution (students, postdocs, and professors):

I thought you may find the following two short comments informative and interesting (both are posted on the Internet). 

1. Molecular clock at best explains half the story on ‘genetic equidistance’ and at worst explains none.

http://thegoldengnomon.blogspot.com/2009_04_01_archive.html

http://precedings.nature.com/documents/1751/version/2

Summary:

The molecular clock is widely known to be problematic but has yet to be put to rest by a knock out punch.  Here, a newly appreciated feature of the original result that provoked the clock and remains the only ‘evidence’ for it was used to ring the death knell for the clock.  The clock is a completely mistaken explanation for the equidistance result (sister species are equidistant to a simpler outgroup) first found by Margoliash in 1963 and should never have been invented in the first place if people had fully understood the equidistance result.  The result has two features.  One is equidistance in terms of percent identity, which originally provoked the clock and remains the only ‘evidence’ for it.  The other is the overlap feature where most of the mutant positions relative to the outgroup are shared between the two sister lineages.  This feature has been completely overlooked for the past 46 years and fully contradicts the clock/neutral theory.  The correct and complete explanation for the equidistance result is the MGD hypothesis that I recently proposed.  Since the clock is totally invalid, results based on the clock are automatically invalid, including the human-chimp split of 5 million years and other important molecular dating reported so far that contradict the fossil record.

 

2.  Convergent evolution, rather than common ancestry, explains the sequence similarity between human and chimpanzees.

http://precedings.nature.com/documents/2123/version/1#comments

http://thegoldengnomon.blogspot.com/

Summary:

It discusses one of the best molecular facts (newly reported in Nature this year) that simply cannot be reconciled in any way with the sister grouping of humans and chimpanzees but fully supports the sister grouping of humans and pongids.  My new molecular dating based on the MGD hypothesis gave a human-pongid spilt time of 19.2 million years, in full agreement with the fossil record.

Would appreciate your feedbacks,

Cheers,

Shi Huang

cc

John Hawks, Milford Wolpoff, Jeffrey Schwartz, Elwyn Simons, David Pilbeam, Michel Brunet, Gen Suwa, Morris Goodman, Christian Schwabe, Laura Katz, Gunter Wagner, Eugene Koonin, Phil Skell, Jerry Harris, Blair Hedges, David Baum, Walter Fitch, Joe Daniel, Sudhir Kumar, Leigh van Valen, James Cai, Laurence Hurst, Tobias Warnecke, David Lambert, Jason Stajich 

The overlap feature of hemoglobin

Hemoglobin was used in 1962 by Zuckerkandl and Pauling to derive data that led to the molecular clock idea.  Here I examined the overlap feature using human Hemoglobin alpha 1 (AAK61216, hba1) to compare with horse and chicken, in a way similar to what Zuckerkandl and Pauling had done.

Hs, Human [Homo sapiens]

Ec, Horse [Equus caballus]

Gg, Chicken [Gallus gallus]

Hs     1    MVLSPADKTNVKAAWGKVGAHAGEYGAEALERMFLSFPTTKTYFPHFDLSHGSAQVKGHG  60

Hs=Ec       MVLS ADKTNVKAAW KVG HAGEYGAEALERMFL FPTTKTYFPHFDLSHGSAQVK HG

Ec     1    MVLSAADKTNVKAAWSKVGGHAGEYGAEALERMFLGFPTTKTYFPHFDLSHGSAQVKAHG  60

Hs=Gg       MVLS ADK NVK  + K+  HA EYGAE LERMF ++P TKTYFPHFDLSHGSAQ+KGHG

Gg     1    MVLSAADKNNVKGIFTKIAGHAEEYGAETLERMFTTYPPTKTYFPHFDLSHGSAQIKGHG  60

Overlap         x          x   x               x 

Non-overlap                                                          *           


Hs     61   KKVADALTNAVAHVDDMPNALSALSDLHAHKLRVDPVNFKLLSHCLLVTLAAHLPAEFTP  120

Hs=Ec       KKV DALT AV H+DD+P ALS LSDLHAHKLRVDPVNFKLLSHCLL TLA HLP +FTP

Ec     61   KKVGDALTLAVGHLDDLPGALSNLSDLHAHKLRVDPVNFKLLSHCLLSTLAVHLPNDFTP  120

Hs=Gg       KKV  AL  A  H+DD+   LS LSDLHAHKLRVDPVNFKLL  C LV +A H PA  TP

Gg     61   KKVVAALIEAANHIDDIAGTLSKLSDLHAHKLRVDPVNFKLLGQCFLVVVAIHHPAALTP  120

Overlap        x    x  x x  x x   x                            x    x

Non-overlap                                                *       *           

 

Hs     121  AVHASLDKFLASVSTVLTSKYR  142

Hs=Ec       AVHASLDKFL+SVSTVLTSKYR

Ec     121  AVHASLDKFLSSVSTVLTSKYR  142

Hs=Gg        VHASLDKFL +V TVLT+KYR

Gg     121  EVHASLDKFLCAVGTVLTAKYR  142

Overlap               x

 

Results:

Of 17 variants between human and horse, 14 are also variants between human and chicken.  So there are 14 overlaps and 3 non-overlaps. 

Molecular clock prediction:  17/142 x 42/142 x 142 = 5 overlap residues, accounting for 36% of total overlap.  If we generously grant 50% positions as absolutely non-variable, we have 17/71 x 42/71 x 71 = 10 overlap residues, accounting for only 71% of total overlap.  To account for 14 overlaps, we need 91 absolutely non-variable residues or require that there are only 51 residues that can vary.  But we know that there are at least 71 variable positions between human and fish.  So, molecular clock simply cannot account for the 14 overlaps.  But the existence of significant overlap is a prime prediction of the MGD hypothesis.


In fact, there are only 21 residues that are absolutely conserved among human, bony fish, lungfish, coelacanths, and sharks.  So, a most realistic calculation of overlap should be 17/121x42/121x122=5.9 residues, far short of 14. 

BTW, chicken is equidistant (70% identity) to human and horse as shown below.  Of 17 variants between human and horse, 14 are also variant between horse and chicken, a significant overlap.

Human-Horse [Equus caballus]:

Identities = 125/142 (88%)

Hs     1    MVLSPADKTNVKAAWGKVGAHAGEYGAEALERMFLSFPTTKTYFPHFDLSHGSAQVKGHG  60

            MVLS ADKTNVKAAW KVG HAGEYGAEALERMFL FPTTKTYFPHFDLSHGSAQVK HG

Ec     1    MVLSAADKTNVKAAWSKVGGHAGEYGAEALERMFLGFPTTKTYFPHFDLSHGSAQVKAHG  60

 

Hs     61   KKVADALTNAVAHVDDMPNALSALSDLHAHKLRVDPVNFKLLSHCLLVTLAAHLPAEFTP  120

            KKV DALT AV H+DD+P ALS LSDLHAHKLRVDPVNFKLLSHCLL TLA HLP +FTP

Ec     61   KKVGDALTLAVGHLDDLPGALSNLSDLHAHKLRVDPVNFKLLSHCLLSTLAVHLPNDFTP  120

 

Hs     121  AVHASLDKFLASVSTVLTSKYR  142

            AVHASLDKFL+SVSTVLTSKYR

Ec     121  AVHASLDKFLSSVSTVLTSKYR  142

 

Human-Chicken [Gallus gallus]

Identities = 100/142 (70%)

Hs     1    MVLSPADKTNVKAAWGKVGAHAGEYGAEALERMFLSFPTTKTYFPHFDLSHGSAQVKGHG  60

            MVLS ADK NVK  + K+  HA EYGAE LERMF ++P TKTYFPHFDLSHGSAQ+KGHG

Gg     1    MVLSAADKNNVKGIFTKIAGHAEEYGAETLERMFTTYPPTKTYFPHFDLSHGSAQIKGHG  60

 

Hs     61   KKVADALTNAVAHVDDMPNALSALSDLHAHKLRVDPVNFKLLSHCLLVTLAAHLPAEFTP  120

            KKV  AL  A  H+DD+   LS LSDLHAHKLRVDPVNFKLL  C LV +A H PA  TP

Gg     61   KKVVAALIEAANHIDDIAGTLSKLSDLHAHKLRVDPVNFKLLGQCFLVVVAIHHPAALTP  120

 

Hs     121  AVHASLDKFLASVSTVLTSKYR  142

             VHASLDKFL +V TVLT+KYR

Gg     121  EVHASLDKFLCAVGTVLTAKYR  142

 

Horse-Chicken [Gallus gallus]

Identities = 100/142 (70%),

Ec     1    MVLSAADKTNVKAAWSKVGGHAGEYGAEALERMFLGFPTTKTYFPHFDLSHGSAQVKAHG  60

            MVLSAADK NVK  ++K+ GHA EYGAE LERMF  +P TKTYFPHFDLSHGSAQ+K HG

Gg     1    MVLSAADKNNVKGIFTKIAGHAEEYGAETLERMFTTYPPTKTYFPHFDLSHGSAQIKGHG  60

 

Ec     61   KKVGDALTLAVGHLDDLPGALSNLSDLHAHKLRVDPVNFKLLSHCLLSTLAVHLPNDFTP  120

            KKV  AL  A  H+DD+ G LS LSDLHAHKLRVDPVNFKLL  C L  +A+H P   TP

Gg     61   KKVVAALIEAANHIDDIAGTLSKLSDLHAHKLRVDPVNFKLLGQCFLVVVAIHHPAALTP  120

 

Ec     121  AVHASLDKFLSSVSTVLTSKYR  142

             VHASLDKFL +V TVLT+KYR

Gg     121  EVHASLDKFLCAVGTVLTAKYR  142

The molecular clock should never have been invented in the first place for macroevolution

Two kinds of sequence alignment can be made using the same set of sequence data.  The first aligns a recently evolved organism such as a mammal against those simpler or less complex species that evolved earlier such as amphibians and fishes.  The second aligns a simpler outgroup organism such as fishes against those more complex sister species that appeared later such as amphibians and mammals.  The first alignment indicates a near linear correlation between genetic distance and time of divergence, implying indirectly a constant mutation rate among different species.  The second alignment shows the genetic equidistance result where sister species are approximately equidistant to the simpler outgroup. This directly triggered the idea of constant mutation rate among different species.  Since both alignments use the same sequence data set, certain information may be revealed by either alone.  But the data that most directly and obviously support the interpretation of a constant mutation rate is the genetic equidistance result. 

The molecular clock hypothesis was first informally proposed by Zuckerkandl and Pauling in 1962 based largely on data from the first alignment [1].  Margoliash in 1963 performed both alignments and made a formal statement of the molecular clock after noticing the genetic equidistance result [2, 3].  “It appears that the number of residue differences between cytochrome c of any two species is mostly conditioned by the time elapsed since the lines of evolution leading to these two species originally diverged. If this is correct, the cytochrome c of all mammals should be equally different from the cytochrome c of all birds.  Since fish diverges from the main stem of vertebrate evolution earlier than ether birds or mammals, the cytochrome c of both mammals and birds should be equally different from the cytochrome c of fish.  Similarly, all vertebrate cytochrome c should be equally different from the yeast protein.”

The results of both alignments have two features.  One is obvious: distance in terms of percent identity, which directly provoked the clock idea.  The other is the overlap feature.  In the post of April 30th, 2009, I explained the overlap feature of the genetic equidistance result.  Here, I show that the first kind of alignment performed by Zuckerkandl and Pauling also shows the overlap feature, as would be expected since both alignments use the same sequence information and should tell the same story.  The clock idea should never have been invented in the first place if Zuckerkandl and Pauling had paid attention to this feature.

 

Again, I use cytochrome c of yeast (Sc), drosophila (Dm), and human (Hs) as an example.  What Zuckerkandl and Pauling had found, when applied in our cytochrome c case here, is that human is closer to drosophila than to yeast.  Human differs from drosophila in 22 positions and from yeast in 36 positions.  The overlap feature in this case is that most of 22 variant positions between human and drosophila are also variant between human and yeast.  This can be easily illustrated in the following alignment:

Dm              -GDVEKGKKLFVQRCAQCHTVEAGGKHKVGPNLHGLIGRKTGQAAGFAYTDANKA

Hs              -GDVEKGKKIFIMKCSQCHTVEKGGKHKTGPNLHGLFGRKTGQAPGYSYTAANKN

Sc              -GSAKKGATLFKTRCLQCHTVEKGGPHKVGPNLHGIFGRHSGQAEGYSYTDANIK

                 *..:** .:*  :* ****** ** **.******::**::*** *::** ** 

 

Dm              KGITWNEDTLFEYLENPKKYIPGTKMIFAGLKKPNERGDLIAYLKSAT

Hs              KGIIWGEDTLMEYLENPKKYIPGTKMIFVGIKKKEERADLIAYLKKAT

Sc              KNVLWDENNMSEYLTNPKKYIPGTKMAFGGLKKEKDRNDLITYLKKAT

                *.: *.*:.: *** *********** * *:** ::* ***:***.**  

The result shows that 17 of the 22 are also variant between human and yeast (these 17 positions are colored in purple and green).  The fact that the overlap is not 100% is because residues conserved due to common adaptation to environment between human and drosophila are different from those between human and yeast. 

The molecular clock predicts: 

The chance for a position to be different between human and yeast is 36/102.

The chance for a position to be different between human and drosophila is 22/102.

The number of overlap positions: 36/102 x 22/102 x 102 = 7.76. far short of 17.



There are only 7 positions as underlined below that are absolutely conserved among bacteria, yeast, plants, nematodes, and human.

Human  1    MGDVEKGKKIFIMKCSQCHTVEKGGKHKTGPNLHGLFGRKTGQAPGYSYTAANKNKGIIW  60

       61   GEDTLMEYLENPKKYIPGTKMIFVGIKKKEERADLIAYLKKAT  103

So, a most realistic calculation of overlap should be 36/95 x 22/95 x 95 = 8.3 residues, far short of 17.

 

Even if we generously grant that 40 residues are absolutely non-neutral or non-variable, we still only get 36/62 x 22/62 x 62 = 12.77, short of 17. 


Again, Zuckerkandl, Pauling, and Margoliash all could have noticed the overlap feature.  If they had done that 46 years ago, the molecular clock (vastly different species have very similar mutation rates) would never have been invented in the first place for macroevolution.  It may have been invented for studying microevolution (identical or very similar species have very similar mutation rates) and may still apply in some cases of microevolution.  But its impact on the understanding of molecular evolution would be trivial. 

Acknowledgements:

I thank my college classmate Dr. Wei Shen for providing the alignment picture shown here, and for many valuable discussions. 

 

Reference:

1.         Zuckerkandl E, Pauling L: Molecular disease, evolution, and genetic heterogeneity, Horizons in Biochemistry. New York: Academic Press; 1962.

2.         Margoliash E: Primary structure and evolution of cytochrome c. Proc Natl Acad Sci 1963, 50:672-679.

3.         Kumar S: Molecular clocks: four decades of evolution. Nat Rev Genet 2005, 6:654-662.


Tuesday, May 5, 2009

Evidence for an ancient adaptive episode of convergent molecular evolution

Castoe et al., “Evidence for an ancient adaptive episode of convergent molecular evolution”  PNAS, April 29, 2009, published on line, doi: 10.1073/pnas.0900233106

 

Abstract: …..These results indicate that nonneutral convergent molecular evolution in mitochondria can occur at a scale and intensity far beyond what has been documented previously, and they highlight the vulnerability of standard phylogenetic methods to the presence of nonneutral convergent sequence evolution.

 

I left a comment on the preprint version of this paper at Nature precedings. http://precedings.nature.com/documents/2123/version/1#comments

 

Indeed, convergent evolution is extremely common. The best illustration of this is a phenomenon I termed ‘genetic nonequidistance to a more complex outgroup’. Thus, relative to a complex outgroup such as human, some sister species from a simple clade are not equidistant to human. The more complex sister species is always closer to human than the simpler sister species. In all five cases (except plants) examined where difference in complexity of the sister species can be inferred (octopus vs. cockle, Terebratulina vs. Lingula, bird vs. snake, dragonfly vs. louse, and smut vs. yeast), the more complex species always show greater sequence similarity to humans.

Because these sister species are separated from humans for the same amount of time, their different sequence similarity to humans must be due to convergent evolution. Thus, sequence similarity to complex species or humans cannot be used to infer closer genealogy with humans.

The sister grouping of chimpanzees and humans really has no other non-ambiguous support other than sequence similarity as measured by percent identity. The premise for this approach has now been nullified by the phenomenon of genetic non-equidistance to a more complex outgroup despite equidistance in time or genealogy. The same premise for grouping an ape (chimpanzee) with human to the exclusion of another ape (orangutan) would equally justify the obviously absurd grouping of human with a mollusk (octopus) to the exclusion of another mollusk (cockle), or with a brachiopod (Terebratulina) to the exclusion of another brachiopod (Lingula), or with a reptile (bird) to the exclusion of another reptile (snake).

The molecular clock hypothesis, i.e., vastly different species have very similar mutation rates, is a tautological interpretation of the ‘genetic equidistance’ result. It is falsified by the ‘genetic nonequidistance’ phenomenon as discussed above. I have recently come up with the ‘maximum genetic diversity’ (MGD) hypothesis to explain equally well both the equidistance and the nonequidistance phenomenon. See my paper posted here, “Inverse relationship between genetic diversity and epigenetic complexity”.

Below is a paragraph from one of my recent manuscripts discussing one of the best facts (newly reported in Nature this year) that simply cannot be reconciled in any way with the sister grouping of humans and chimpanzees but fully supports the MGD hypothesis and the sister grouping of humans and pongids.

Consistent with low genetic diversity in humans, human specific segmented duplications show lower copy number polymorphisms in humans than chimpanzee specific segmented duplications do in chimpanzees [54]. Similarly, those duplications shared among human, chimpanzees, and orangutans, or those shared among human, chimpanzees, orangutans, and monkeys are also less polymorphic in humans than in chimpanzees, indicating clearly that duplications that are shared because of common ancestry are less polymorphic in humans than in chimpanzees. In contrast, the duplications shared between human and chimpanzees are equally polymorphic in humans and chimpanzees. This unusual result contradicts the sister grouping of humans and chimpanzees, because both the MGD and the bottleneck hypothesis would predict lower polymorphism in humans if these duplications are shared because of common ancestry. However, it is fully consistent with the interpretation that the shared duplications between humans and chimpanzees are not due to common ancestry but are due to common selection of independent duplications. Common selection leading to shared sequences is well established [55]. The MGD hypothesis interprets many of the shared sequences between human and chimpanzees as a result of common selection rather than common ancestry. The similar selection pressure leads to similar levels of polymorphism. This result is thus one of the best that simply cannot be reconciled in any way with the sister grouping of humans and chimpanzees but fully supports the MGD hypothesis and the sister grouping of humans and pongids.

Ref:
54. Marques-Bonet T, Kidd JM, Ventura M, Graves TA, Cheng Z, et al. (2009) A burst of segmental duplications in the genome of the African great ape ancestor. Nature 457: 877-881.
http://www.publicacions.ub.es/refs/micoshumans.pdf

55. Bull JJ, Badgett MR, Wichman HA, Huelsenbeck JP, Hillis DM, et al. (1997) Exceptional convergent evolution in a virus. Genetics 147: 1497-1507.

Genomics, Evolution, and Pseudoscience: T. rex protein degrades further

Genomics, Evolution, and Pseudoscience: T. rex protein degrades further

Monday, May 4, 2009

Molecular clock, pseudoscience, and Steven Salzberg

I have left several comments before on Professor Steven Salzberg’s blog:

http://genome.fieldofscience.com/2008/08/t-rex-protein-degrades-further.html

He is a professional in the business of molecular clock and molecular evolution.  But he has a hobby of carelessly calling the research of some scientists ‘pseudoscience’.  (granted that he may have made some good calls occasionally)  I was really surprised to see his aggressive and baseless attack on the dinosaur peptide work of John Asara et al, which I have made use in my paper on testing the molecular clock using fossil sequences. His blog title is “Genomics, evolution, and pseudoscience”.  So, given his ‘expertise’ on both molecular evolution and pseudoscience, I sent a post to his blog the other day to show him why his field, the molecular clock, may be pseudoscience.  But my post has yet to appear on his blog after more than 2 days (he screens all post before posting them and my posts have all went through prior to this latest one).  Most likely, he would not post it.  If he can only be silent to my analysis, it could only mean that he cannot refute it (he would have to be extremely stupid to try to refute it, because it is irrefutable).  If he is a genuine scientist, he would post it regardless whether it is true or false.  So here we have a good joke, an active practitioner of pseudoscience makes it a hobby exposing pseudoscience except his own. I hope he can sleep in peace now that he knows there is at least one observer who knows what a fake he is.  Good luck to him.

 

Below is what I sent to his blog:

 Steven,

I submit the following analysis to suggest that the molecular clock is a candidate for the most outrageous pseudoscience in the recent history of science.  As an expert on both pseudoscience and molecular evolution, you are well qualified to render an honest and impartial evaluation of this analysis.

 Molecular clock at best explains half the story on ‘genetic equidistance’ and at worst explains none. 

(detail omitted here as it is the same as the post of April 30, 2009 on my blog)

Friday, May 1, 2009

More on the MGD interpretation of certain facts

 

I am showing below yeast and neurospora cytochrome c alignment to explain an observation that is not covered in detail in my MGD paper but is covered in principle.  The observation is that not all variant residues between two complex species are also variant between two simple species, when Table 1 of the paper seems to indicate that all variant residues between two complex species are also variant between two simple species.  Table 1 is of course meant to express an idea in simplistic form and should not be taken to be literally exact. 

 

 Identities = 71/105 (67%), Positives = 88/105 (83%), Gaps = 0/105 (0%)

 

yeast  3    FKAGSAKKGATLFKTRCLQCHTVEKGGPHKVGPNLHGIFGRHSGQAEGYSYTDANIKKNV  62

            F AG +KKGA LFKTRC QCHT+E+GG +K+GP LHG+FGR +G  +GY+YTDAN +K +

Neuro  3    FSAGDSKKGANLFKTRCAQCHTLEEGGGNKIGPALHGLFGRKTGSVDGYAYTDANKQKGI  62

 

yeast  63   LWDENNMSEYLTNPAKYIPGTAMAFGGLKKEKDRNDLITYLKKAT  107

             WDEN + EYL NP KYIPGT MAFGGLKK+KDRND+IT++K+AT

neuro  63   TWDENTLFEYLENPKKYIPGTKMAFGGLKKDKDRNDIITFMKEAT  107

 

 

There are 10 of 22 residues differing between drosophila and human that are also different between yeast and neurospora. 

 

So a major portion (about 50%) of variants between drosophila and human is also variant between yeast and neurospora.  That portion can only be explained by the MGD but not by the molecular clock/neutral theory.  Now, why not all? 

 

As shown by Figure 2 of my paper, the MGD says that of all the conserved residues between two species at any time, there are a fraction of them that is due to adaptation to common environmental selection and may change from time to time.  In our case, yeast and neurospora share 67% of all positions.  Of these, maybe 20% is shared because of common selection (the two yeasts have very similar way of life).  So, the absolutely non-neutral sequence may be 47% and the neutral region 53%.  Human and drosophila share 80% positions, and 10% of these may be due to common environmental selection (they have very different life style and so the shared region due to common selection is less).  So the neutral sequence is 30% in this case for the drosophila.  The MGD says that this 30% region should completely overlap the 53% neutral region of yeast.  But since only 20 of 30 in human/drosophila did actually vary and only 30 of 53 in yeast, the actual overlap residues are 20/30 x 30/53 x 30 = 11, which is very close to the actual number. 

 

So, the exact numbers may not be real but the above is to illustrate in principle how the observation may be explained by the MGD.  The key is to have a fraction of the shared residues as being neutral or changeable with environment, even when the distance is at maximum.  This is a very reasonable and intuitively obvious point and is actually true in reality.  So, when we see a maximum distance, it does not mean that all the shared residues are absolutely nonchangeable or nonneutral.