Showing posts with label infinite sites. Show all posts
Showing posts with label infinite sites. Show all posts

Saturday, September 5, 2015

The branch of biology with the most mathematics is also widely known as the most soft! Why?

The field of population genetics and molecular evolution was largely founded by mathematicians/statisticians such as Fisher, Haldene, and Wright. Even pure mathematician like Hardy has contributed a key equation to the field. But contrary to naive expectations, this field of study is more like soft social science than to hard core physics. In the words of Jerry Coyne (an extremely enthusiastic propagator of the Darwinian evolution theory and a professor of evolutionary studies at the University of Chicago): "In science's pecking order, evolutionary biology lurks somewhere near the bottom, far closer to phrenology than to physics".

Why? The reason is simple. Math depends on assumptions or paradigms. The assumptions for hard core sciences are axioms or self evident intuitions. Euclid and Newtons axioms come to mind. They are all a priori true or self evidently true. In contrast, there is not a single assumption in the evolution field that is self evidently true or can qualify as axiom. Nearly all assumptions in that field are in fact self evidently false. Just a few examples, the infinite sites model, the neutral/junk DNA assumption, random mating, and the independent mutations assumptions.

Key figures in the field has also acknowledged this, as Ohta and Gillespie said in 1996: "all current theoretical models suffer either from assumptions that are not quite realistic or from an inability to account readily for all phenomena." (Theoretical Population Biology,1996, 49: 128-142) 

To show a flavor of the amount of math in the field, below are two pages of my notebook from an undergrad evolutionary genetics course taken 32 years ago at my Alma mater Fudan University.



Tuesday, July 28, 2015

Begging the question, a common practice in evolutionary genetics (junks assumed, then deduced)

The molecular evolution and popgen field is known to have the most mathematics among all branches of biology. But precisely because of that, it needs many simplifying assumptions or premises, which often lead to the fallacy of begging the question. I here gave a few examples regarding the concept of the mutation or genetic load and the genetic load argument for junk DNA, based on reading a paper on the genetic load (Lesecque et al 2012). Some have used the genetic load as the best argument for the junk DNA notion (Palazzo and Gregory, 2014; Graur, 2015). I also discuss the assumption of non-conservation equaling non function and the assumption of infinite sites.

Lesecque et al say: “The mutation load was more formally defined as the proportional reduction in mean fitness of a population relative to that of a mutation-free genotype, brought about by deleterious mutations (Crow 1970):

L = (Wmax - Wmean)/Wmax

where Wmean is the mean fitness of the population at equilibrium and Wmax is the mean fitness of a deleterious mutation-free individual.”

Is there a deleterious mutation-free individual in a real world or even an imagined world? All mutations, as random mistakes, have a deleterious aspect to an ordered system, if not individually, then collectively. Many mutations could be both deleterious and beneficial. For example, they could be beneficial to adaptive immunity that requires genome variation for producing diverse antibody responses but deleterious to innate immunity that requires conserved proteins to recognize conserved sequences shared by a certain class of microorganisms. By failing to recognize the both deleterious and beneficial nature of most mutations and by classifying mutations into two kinds (deleterious and non-deleterious with the latter consisted of mostly neutral ones), the assumption on the concept of deleterious and non-deleterious mutations eventually led to the genetic load argument for the conclusion that most mutations must be neutral. Here, one sees that the neutral conclusion is already embedded in the premise that led to it. The premise does not recognize the fact that most mutations appear neutral or nearly neutral as a result of balancing selection, and the fact that all mutations have a deleterious aspect as noises to a finely tuned system. Of course, that premise works for a junkyard like system.

Lesecque et al say: ““If the fitness effects of deleterious mutations are independent from one another, the mutation load across all loci subject to recurrent mutation is approximately

L = 1-e-U

(Kimura and Maruyama 1966), where U is the overall rate of deleterious mutation per diploid genome per generation. This simple formula is a classic result of evolutionary genetics.”

So, a classic formula for the genetic load argument is based on the assumption that the fitness effects of deleterious mutations are independent from one another. For a junk yard, yes, the consequences of errors in the building parts are independent from one another. However, for a system that is ordered and built by network-like interactions among the building parts, no, the consequences of errors in the building parts are NOT independent from one another. In fact, recent studies in genomics are constantly discovering epistatic interactions among mutations. So, here one sees clearly again, the neutral or junk DNA conclusion is already embedded in the premise that treats an organism more as a junkyard than a highly ordered system with components organized in a network fashion. When you have already assumed an organism to be junk like, why bother showing us the math formula and deduction leading to the junk DNA conclusion? You should just say that most DNAs are junks because I said so.

Finally, none of the premises related to the genetic load concept recognized the fact that a large collection of otherwise harmless mutations within an individual could be deleterious, as our recent papers have shown. Well, again, such a fact certainly does not exist for a junkyard-like system. By not recognizing that fact or being too naïve to see it, the practitioners in the popgen field have again and again assumed biological systems to be junk like before setting out to prove/deduce that they are made of largely junks.

I also briefly comment on a paper by the Ponting group concluding that human genome is only about 8% functional (Rands et al, 2014). The premise for that deduction is that non-conservation means non-function. Again, building parts for different junk yards are not conserved and nonfunctional. So, non-conservation means non function holds for junk yards. But for organisms relying on mutations to adapt to fast changing environments, recurrent or repeated mutations at the same sites at different time points in their life history are absolutely essential for their survival. Less conserved sequences are more important for adaptation to external environment, while the more conserved ones are important for internal integrity of a system. For bacteria or flu viruses to escape human immunity or medicines, the fast changing or non-conserved parts of their genome are absolutely essential. So, here again, by assuming non-function for the non-conserved parts of the genome, one is assuming an organism to be like a junk yard.

Other key assumptions like the infinite sites model (means neutral sites) are critical for phylogenetics as it is practiced today and for the absurd Out of Africa model of human origin that uses imagined bottlenecks to explain away the extremely low genetic diversity of humans. Well, a junk yard can certainly have an infinite number of parts and tolerate an infinite number of errors. An organism’s genome is finite in size and essentially nothing compared to infinite size. Within such finite size genomes, the proportion that can be free to change without consequences is even more limited or finite.

A paradigm shift (or revolutionary science) is, according to Thomas Kuhn, a change in the basic assumptions, or paradigms, within the ruling theory of science. The above analyses show that the assumptions for the popgen and molecular evolution field are largely out of touch with reality as more reality becomes known, and must be changed quickly if the field wants to avoid fading into oblivion and stay relevant to mainstream bench biology, genomic medicine, archeology, and paleontology. Those assumptions have produced few useful and definitive deductions that can be independently verified and avoid the fate of constant and endless revisions, like we have seen from 1987 to now for the Out of Africa model or the Neanderthals.

Lesecque Y, Keightley PD, Eyre-Walker A (2012) A resolution of the mutation load paradox in humans. Genetics 191: 1321–1330 .

Palazzo AF, Gregory TR (2014) The Case for Junk DNA. PLoS Genet 10(5): e1004351. doi:10.1371/journal.pgen.1004351

Dan Graur (2015) If @ENCODE_NIH is right each of us should have on average from 3 × 10^19 to 5 × 10^35 children. https://www.dropbox.com/s/4bj3andtlu3y9hk/Genetic%20mutational%20load.docx?dl=0 …


Rands CM, Meader S, Ponting CP, Lunter G (2014) 8.2% of the Human Genome Is Constrained: Variation in Rates of Turnover across Functional Element Classes in the Human Lineage. PLoS Genet 10(7): e1004525.


Monday, October 27, 2014

Why the surprising pattern of no genetic continuity between people living in the same area but from different periods of time? Think the flu virus!

I used three slides as shown below to illustrate the idea of informative DNAs in my talk in last month’s workshop on genome and evolution in Naples, Italy.

The antigenic sites in human influenza A virus mutate and turn over quickly, which is critical for their survival or escape from human neutralizing antibodies and hence responsible for flu epidemics. As shown in Figure 1, two amino acid positions in hemagglutinin (156 and 145, panel a and b) turned over several times within a 30 year period, while two others (138 and 194, panel c and d) stayed largely unchanged (Figure from Shih et al, 2007). 

The flu results illustrate two important points with regard to evolutionary dynamics of a genome that have so far been grossly overlooked by the evolution and popgen field. First, fast evolving or less conserved DNAs are also functional rather than neutral as they are essential for quick adaptive needs in response to fast changing environments. Second, fast evolving DNAs turn over quickly and can be shown to violate the infinite sites model.  Hence, they cannot be used for phylogenetic inference. If one uses the fast changing sites in a flu virus to infer the phylogenetic relationship of the virus isolates responsible for different epidemics in a past period of say 10 years, one would reach the absurd conclusion that each epidemic was caused by a distinct type of flu virus with no genetic continuity among them rather than just minor variations of the same type.

Mutation rates in humans are of course much slower than that in a flu virus. But just like a flu virus, there are also fast and slow changing sites (Figure 2). The time scales are different but the principle is the same.  The fast changing sites may turn over every few thousand years and in fact make up the majority of the observed variant sites in humans when properly examined by us (Figure 3). This is why the field of ancient DNA kept producing the absurd pattern of no genetic continuity between people living in the same area but from different periods of time. All of the published analyses have simply used the wrong sites that are equivalent to the fast changing antigenic sites in a flu virus. What one should be using are sites with very slow mutation rates, like 1 mutation every 50,000 years. We have been busy reinterpreting the published DNAs for several years now and hope to submit our work soon.


Figure 1. (a and b) Frequency changes at residue sites 156 (a) and 145 (b) were highly dynamic. (c and d) Sites 138 (c) and 194 (d) did not undergo major frequency change over time.




Figure 2. A priori model of evolutionary dynamics of human genomic DNAs.




Figure 3. Difference between slow and fast evolving sites. Shown are a piece of homologous DNA in three different individuals or species. In the fast evolving DNAs making up the vast majority of human genome, there is obvious and verifiable violation of the infinite sites model. These DNAs have abundant overlapped mutant sites where independent mutations have occurred on the same site in different individuals or species. 


Ref.


Shih, C-C., Hsiao, T-C., Ho, M-S., and Li, W-H. (2007) Simultaneous amino acid substitutions at antigenic sites drive influenza A hemagglutinin evolution. Proc Natl Acad Sci U S A. 104:6283-6288.

Sunday, August 17, 2014

Testing the infinite sites assumption

We presented the following poster at the "1000 genomes and beyond" meeting, Cambridge, UK, 24-26, June 2014. We are also going to present it in the ASHG 2014 meeting in Oct 2014, San Diego. 

The bottom line is that there are very few neutral or junk DNAs in the human genome, at least when one examines the genome by using experimental approaches. (a new paper of ours on disproving the neutral theory by using an experimental approach has just got published here, titled Scoring the collective effects of SNPs: associations of minor alleles with complex traits in model organisms.)  All previous studies used only bioinformatics approaches. Their conclusions of less than 10% functional genome are based on UNCERTAIN assumptions and therefore are mostly meaningless. The field must realize that it is time to stop such senseless researches based on senseless assumptions. We should either do experiments without any prior assumptions or if we have to, we must only use a priori sound intuitions as our assumptions.

The abstract, introduction and discussion of the our poster are posted here. The poster can be downloaded from my lab website.

  Abstract The infinite sites model of the neutral theory is a fundamental assumption underlying nearly all population genetic and phylogenetic studies today but has yet to be properly tested. We here tested it from two novel perspectives using the 1000 genomes dataset. First, we examined the genetic diversity patterns of different human populations using a variety of different types of SNPs, such as a random set of SNPs representing genome average, stop codon, nonsyn, syn, etc. Patterns shown by a random set of SNPs are expected to be similar to those shown by known functional stop codon SNPs, if most SNPs are not neutral. In contrast, neutral SNPs should show a most different pattern from stop codon SNPs. Second, it has long been well known that most genetic variations are shared among different human groups, which has been interpreted from the infinites sites perspective to mean few genetic differences among the ethnic groups (Lewontin, 1972). But the possibility of saturation or independent mutations to account for this phenomenon has yet to be examined and excluded. We compared the number of shared SNPs in DNAs of different evolutionary rates among different human populations to see if shared SNPs are in fact a result of independent mutations or saturation and hence more common in fast evolving DNAs relative to slow ones. We found that a random set of SNPs are just like the stop codon SNPs in showing Africans to have the largest genetic diversity. Shared SNPs are enriched in fast evolving DNAs. These results suggest that the vast majority of the human genome do not follow the infinite sites model.

     Introduction  Molecular studies have so far relied on the Neutral theory and its infinite sites assumption. The Neutral theory was originally inspired by the so called molecular clock which was in turn inspired by the first and most remarkable result in molecular evolution, the genetic equidistance result that sister species are approximately equidistant to a simpler outgroup. In recent papers, we have shown that the equidistance result has been incorrectly interpreted by the molecular clock with grave consequences on phylogenetic studies: nearly all past studies have used non-informative DNAs assumed to be neutral but have now been shown by us to be under selection (Hu et al, 2013, Huang, 2010). The neutral theory was mistaken right from its inception. We have developed the maximum genetic diversity (MGD) hypothesis to absorb and supersede the neutral theory (Hu et al, 2013). From this more correct/complete theoretical perspective, we here tested whether the infinite sites model holds for the majority of the human genome as is commonly assumed.  

       Discussion: mutation rate, sequence conservation, and neutrality
       The results suggest that the vast majority of human genome do not follow the infinite sites model and are not neutral. Only a very limited sites: the non-syn slow evolving SNPs as defined here, behaved uniquely among all the SNPs examined and appear to be neutral or follow the infinite sites model. They are not deleterious as they are different from stop codon SNPs.  They are also not under positive selection as positively selected genes tend to be fast evolving.  To the dramatic difference between slow and fast evolving DNAs as shown here, we cannot come up with a meaningful explanation using any known schemes other than the recently proposed idea of maximum genetic diversity. 
       Variation in mutation rate in different regions of the human nuclear genome may exceed 1000 fold.  That a gene is slow evolving could be due to at least two reasons.  One is being located in a region of the genome with slow mutation rates.  This however may not apply to the difference in mutation rates between non-synonymous and synonymous sites of the same gene as found here.  Alternatively, most mutations may hit functional sites and be negatively selected by the need to maintain the internal integrity/order of a biological system. It would take many mutations and hence a long time before a neutral site is hit, thus giving the appearance of a slow mutation rate. Since changes in such neutral sites take long time, they may be too slow to meet adaptive needs to be under positive selection.  Given the apparent slow rate and absence of positive selection, they are also unlikely to reach excess levels to cause harm or be under negative selection. 
        Hence, sequence conservation per se may not automatically indicate functionality of variants within such sequences as is commonly assumed.  Less conserved sequences are more important for adaptation to external environment, while the more conserved ones are important for internal integrity of a system. To a virus or bacteria facing elimination by human medicines, the fast evolving parts of their genome is far more critical/functional to their survival than their more conserved parts. The popular assumption of neutrality/non-functionality for the less conserved parts of the genome overlooks their fundamental function in quick adaptation.