Showing posts with label evolution. Show all posts
Showing posts with label evolution. Show all posts

Wednesday, 23 February 2022

ARTICLE: Corals to crops -- how life protects the plans for its cellular power stations

Avoiding organelle mutational meltdown across eukaryotes with or without a germline bottleneck
David M Edwards, Ellen C Røyrvik, Joanna M Chustecki, Konstantinos Giannakis, Robert C Glastad, Arunas L Radzvilavicius, Iain G Johnston
PLoS Biology 19 e3001153 (2021)

(this text is from a press release about the article)

An international team of researchers led by the University of Bergen has uncovered how organisms from crops to corals may avoid deadly DNA damage during evolution.

Our cells, and those of animals, plants and fungi, contain compartments that produce chemical fuel. These compartments contain their own DNA, which stores instructions for important cellular machinery. But this so-called oDNA (organelle DNA) can become mutated, corrupting the instructions and preventing cells making enough energy.

In humans and some other animals, a process called the “bottleneck” allows some offspring to inherit less mutated oDNA. This process needs mothers’ egg cells to develop early, like in humans, where a human girl is born with all her egg cells already formed. But other organisms, from plants to fungi, don’t develop these cells early – their flexible body plans mean that eggs are not “set aside” early in development.<\p>

“We wanted to know how these organisms might avoid inheriting mutations without a human-like bottleneck,” said Ellen Røyrvik, a geneticist on the research team, based at UiB.

The scientists used mathematical modelling to show that a process called gene conversion – the controlled overwriting of DNA – could in theory allow some offspring to inherit less mutant oDNA without requiring a bottleneck. Using genome data, they found machinery controlling this process in plants and fungi, but also in soft corals, sponges, and algae – all organisms without fixed body plans. They also found that this machinery was most active in the parts of plants that will end up producing the seeds of the next generation, suggesting that it is indeed used to allow some offspring to inherit fewer mutations.

Organisms without fixed body plans (including octocorals, sea pens, sponges, plants, and fungi) and with fixed body plans (including humans and many animals) may use different strategies to avoid the buildup of damage in their cellular "power stations." CREDIT: Gemma Lofthouse

“Taken together, it looks like organisms without a fixed body plan – plants, fungi, corals, sponges, algae – may have adopted gene conversion to deal with oDNA mutations,” said Iain Johnston, an associate professor in the Mathematics Institute at UiB, who led the research. “Humans and other animals can develop egg cells early and use a bottleneck; other organisms can use gene conversion instead.”

Going forward, the team plans to explore how this overwriting of oDNA causes other issues in the organisms that use it – including crop plants, where it can cause sterility. They are also exploring the broader question of why these compartments contain oDNA at all, given the risk of mutational damage.

Friday, 23 April 2021

ARTICLE: How does tool use evolve in animals?

Data-Driven Inference Reveals Distinct and Conserved Dynamic Pathways of Tool Use Emergence across Animal Taxa, iScience 23 101245 (2020)

There are some wonderful examples of animals using tools. Octopuses block oyster shells open with coral; boxer crabs wave captive anemones for defense and food capture; captive dolphins use feather to wipe clean their aquarium windows. A while ago, we saw this excellent infographic in National Geographic, and got interested in these data. Do different families of animals learn to use tools in completely different ways? Or are there some general ("universal") principles behind how animals learn to use tools?

There is, of course, a big and fascinating literature on tool use, but we found rather few studies attempting a quantitative comparative analysis across bilaterian animals. To address other evolutionary questions, we've developed HyperTraPS (hypercubic transition path sampling), a statistical approach for learning the "pathways" of evolutionary processes. That is, which events occur before and after which other events in an evolving system? Does feature A always evolve before feature B? We used HyperTraPS to ask about the orderings with which different types of tool use appeared in animals. For example, do animals always learn to "poke" before they learn to "dig"? Do all animals learn tool use in the same order, first A then B then C..., or does it vary across species?

We found some answers that we think are quite interesting. There seem to be some quite deep similarities across animal species in how tool use evolves. Types of tool use like "affixing" and "throwing" are almost universally acquired early; types like "cutting" and "symbolising" are acquired late and rarely, only by primates. The environment and animal family influences the structure of these pathways: aquatic organisms seem to discover "waving" tool use relatively early, for example, and primates discover tools that "block" relatively late.

(A) The inferred pathways of tool use emergence across animals. The size of a blob gives the probability that that mode of tool use (on the horizontal axis) is acquired at that stage (on the vertical axis) of a species' discovery of tool use types. (B) Sample evolutionary pathways of tool use, with individual animal lineages illustrated at the positions corresponding to the modes of tool use they have discovered.

Of course, there's a lot of uncertainty about any analysis like this. Are we talking about wild or captured animals? What if we just haven't observed some types of tool use? We attempted to address several such questions with our analysis and showed that our overall results were quite robust with respect to these uncertainties. HyperTraPS fully describes the uncertainty in its outcomes, helping interpretability. We hope that our results help at least to suggest some possible principles and points for further investigation in this fascinating topic. You can read more in iScience here.

Thursday, 11 July 2019

ARTICLE: Tension and Resolution


Tension and resolution: dynamic, evolving populations of organelle genomes within plant cells
IG Johnston
Molecular Plant 12 764 (2019)


Mitochondria and chloroplasts are compartments in cells that power complex life. Both started out billions of years ago as independent organisms with complete genomes, that were acquired by ancestral cells. Since these endosymbioses, the genomes of mitochondria (mt) and chloroplasts (cp) have become stripped down. Modern mt and cp have lost lots of genes either completely or the “host” cell nucleus. Mt and cp now exist in dynamic populations within the cells of modern organisms. In plants and algae, the two co-exist, sharing responsibility for the energy balance of the organism – and hence ultimately powering and feeding life, including the human population.

Plant mt and cp populations are weird. Different plants and algae have very different mt and cp genomes – some huge (many megabases, several chromosomes in the case of some mt) and some tiny. Unlike the more familiar animal (and human) case, plant mt genomes readily recombine, mixing up their structures and genetic content within the cell. Both mt and cp move around plant cells rapidly – we’re not sure why, particular for mt. Again, unlike animal mt, neither plant mt not cp are particularly prone to meet up and fuse into big networks – they usually stay as individual compartments, except for short interactions. We do know that if we perturb the physical or genetic dynamics of organelles, the plant suffers – which we can sometimes exploit in breeding efficient crops.

 Populations of mitochondria (A green, B) and chloroplasts (A blue, C) moving in the plant cell

In a recent review article here in Molecular Plant, we reviewed current knowledge about these dynamics and speculated about what principles these populations of mt and cp may be responding to. We first asked why mt and cp may retain different sets of genes in different species – a question we’ve touched upon before here (blog). Retaining more genes in organelles may have the “pro” of making individual organelles more independent, and better at responding to demands (see John Allen’s CoRR hypothesis, e.g. here). But there’s the “con” that organelles are dangerous places, and genes retained there may be more subject to damage than in the safe haven of the nucleus. So individual plants may choose to retain mt and cp genes for dynamism, or shift them to the nucleus for robustness. Neither extreme is perfect – there are always pros and cons – leading to a tension to which different plants have selected different resolutions.

Pursuing this line, we next speculated that because plants are immobile (and hence unable to move away from challenging conditions), they may favour the “dynamism” side over the “robustness” side. This would explain why they often retain more organelle genes than motile organisms, but would also predict that they face a double challenge: (i) more organelle genes and (ii) exposure to more challenging environments, both of which may lead to genetic damage. This could be a reason why plant organelles undergo recombination – as a way of ameliorating genetic damage. But again, there are pros and cons: the “pro” of fixing genetic damage is balanced by the “con” of recombination mixing and confusing genetic structure. Perhaps this is why the physical behaviour of plant organelles is different to that in animals – keeping mt and cp separate may limit the amount of recombination that can take place, allowing the plant to control this second pro-con tradeoff.

(left) the proposed tension between robustness (i) and dynamism (ii). Perhaps plants are more (ii)-like because they need to respond to fluctuating conditions... because of their immobility (right) with hypothesised knock-on consequences.

All of these ideas are presented as hypotheses, and we proposed some ways that a combination of new experiment and theory can help make progress understanding these complex, vital systems in future. Watch this space! Iain

Saturday, 22 September 2018

ARTICLE: How do plants roll dice?

Johnston, I.G. and Bassel, G.W. Identification of a bet-hedging network motif generating noise in hormone concentrations and germination propensity in Arabidopsis. Journal of the Royal Society Interface15 141 (2018)

Seeds feed the world, and uniform, reliable harvests of seeds and grains is essential for food security. However, there's a fundamental tension between the evolutionary priorities of plants and the agricultural priorities of humans. Evolutionarily, it is good for plants to "hedge their bets" by having seeds germinate at different times. A plant whose seeds all germinate in March will be susceptible to a frost in April, potentially leading to the loss of a generation of offspring. By contrast, a plant whose seeds germinate throughout March and April will have a subset of its offspring survive that frost, and its genes will be passed on to the next generation.


This bet-hedging poses a challenge for agriculture. In agricultural settings, we have more control over plant environments, and so plants have less need to withstand unpredictable environmental fluctuations. At the same time, non-uniform germination decreases crop yields, makes harvesting harder, and makes crops more susceptible to pest invasion. If we can learn how plants generate this evolved germination variability, we can design engineering and/or breeding strategies to reduce this and improve crop yields.



Plants have evolved to "hedge their bets" by having seeds germinate at different times -- this makes generations of plants more robust to environmental fluctuations. Our work reveals a mechanism that "rolls dice" within plant cells, acting like a random number generator to produce variability in germination propensity. 

In a previous paper (blog post here), we looked at how germination is controlled by an interaction between two hormones known as ABA and GA. During that project, we noticed a surprising feature of the cellular pathways affecting ABA. Oddly, it seemed that ABA both activated a pathway that increased its own production, and at the same time (and in the same place) activated a pathways that increased its own degradation. These two pathways seemed to be competitive -- one increases levels of ABA, the other decreases them. Why would cells spend energy in this "futile" way?


We hypothesised that these competitive pathways might have the effect of generating variability in ABA levels. The pathways are fundamentally "noisy", involving random interactions in the chaotic environment of the cell. Consider increasing the activity of both pathways simultaneously. One pathway would act to increase levels of ABA, the other would act to decrease it. The increased "push and pull" of these noisy pathways would increase the spread of levels of ABA in different cells, even if average levels stayed the same.


Because it's hard to measure the levels of hormones in individual cells over time, we initially took a theoretical approach. We showed, with maths, that the competing pathways did indeed have this variability-inducing effect. By varying the activity through these pathways, the cell can increase variability in ABA levels, and hence increase variability in germination propensity. We showed that the theory we developed was compatible with some experiments where the ABA circuitry was artificially manipulated. The theory went on to reveal various aspects of cellular machinery that we could conceivably target through synthetic approaches, in order to reduce germination variability. Put together, our quantitative theory, supported by experiment, explained the mysterious competitive pathways and revealed several new interventions with the potential to improve food security. You can read about it for free in the Journal of the Royal Society Interface here. Iain  


ARTICLE: Which genes are essential for bacterial survival?

Goodall, E.C., Robinson, A., Johnston, I.G., Jabbari, S., Turner, K.A., Cunningham, A.F., Lund, P.A., Cole, J.A. and Henderson, I.R., 2018. The essential genome of Escherichia coli K-12. mBioe02096 (2018)

Bacteria cause diseases, and are developing resistance to the drugs we use to kill them. Anti-microbial resistance (AMR) is one of the most pressing global health challenges facing society. In the immense scientific endeavour of creating new, effective treatments for bacterial infections, fundamental biological knowledge about how bacteria live and proliferate is of vital importance.


One way we can obtain this knowledge is by discovering what cellular machinery that bacteria need to survive and proliferate. A common (and famous) bacterium called Escherichia coli (E. coli) has over 4000 protein-coding genes, but we're not really sure which of these genes is essential for the bacterium, and how many provide some non-essential "added value". If we can learn which genes are essential for bacteria, we have a more specific set of targets to shoot for in designing new drugs and therapies.


So -- how can we find out which genes are essential for E. coli? One neat way involves a new experimental approach called transposon-directed insertion site sequencing (TraDIS). Transposons are elements of DNA that can be inserted into a bacterial genome -- when they are inserted into part of the genome that codes for a gene, they prevent that gene being properly expressed, effectively removing it from the bacterium. TraDIS, in essence, takes a large population of bacteria and inserts one transposon into a random position in each bacterium. The population is then left to evolve for some time. After that time, we look at the genomes of bacteria within the surviving population, and see exactly where transposon insertions have been retained in some living bacteria.



A stylised representation of the E. coli genome and the positions within it where we found transposons to have been retained (corresponding to non-essential genes). 

The idea is that any bacteria in the population that have a transposon inserted into an essential gene will die. As such a gene is essential, it's required for survival, and a transposon preventing its expression will kill the bacterium. Therefore, if some bacteria in a population retain an insertion in gene X and survive, it follows that gene X is not essential. Conversely, if we see a large region of the genome within which no insertions are retained in the final population, it is likely that that region corresponds to an essential gene. 


There's some mathematical subtlety in the "it is likely". Depending on how many transposon insertions originally occur, and the length of the genome, some regions without insertions may occur just by chance. We did a bit of maths to work out how unlikely it is to see an insertion-free region of a given length arise by chance; and, by extension, how likely it is that a gene identified by this analysis is indeed essential for the bacterium. However, the maths was only one part of this project -- it was first and foremost an experimental tour de force by our excellent collaborators. We jointly provided a new atlas of essential genes in E. coli, provide a new way of reasoning about the powerful TraDIS technique, and provide several new insights into bacterial physiology and biochemistry. The work is freely available in the journal mBio here. Iain 

Saturday, 10 September 2016

ARTICLE: Migration, mothers, mitochondria, and medicine

mtDNA diversity in human populations highlights the merit of haplotype matching in gene therapies

EC Røyrvik, JP Burgstaller, IG Johnston
Molecular Human Reproduction 22 (11), 809-817 (2016)
  • The diversity of mtDNA in modern human populations may pose a challenge to gene therapies that aim to prevent the inheritance of deadly mtDNA disease; we use population and census data, and large-scale mtDNA sequence data, to assess this risk and suggest strategies to combat it.
Some mothers carry disease-causing mutations in their mitochondrial DNA (mtDNA), which can be passed on to their children. Amazing cutting-edge therapies are designed to avoid the inheritance of mutant mtDNA, by endowing a child with mtDNA from another woman (let's say Wilma) -- with no dangerous mutations -- instead of the mother's (let's say Miranda's) mtDNA. However, due to technical challenges in the implementation of these therapies, a small amount of the mother's mtDNA may remain in the child. If that initially small amount can become amplified -- say Miranda's mtDNA proliferates more quickly than Wilma's -- it may come to dominate cells in the child. Then the disease which the therapy attempted to avoid may become manifest -- as we've written about before

We have previously found, in mice, that the more different two mtDNA types are, the more likely one is to dominate over another. So if Miranda and Wilma have very different mtDNA, there's a good chance Miranda's might become amplified. But, although these effects are dramatic in natural mouse populations, we don't really know how likely this "winning" and "losing" was between human mtDNAs (as we'd see in the above therapies). Say Matilda and Wilma both come from London. How different will their mtDNA types likely be? And so, what is the risk that Matilda's mtDNA will beat Wilma's, potentially complicating therapies?


Human mtDNA varies by geography -- women from different parts of the world belong to different mtDNA "haplogroups". Some haplogroups are themselves very diverse, and some less so; haplogroups also differ from each other by varying degrees. So we needed to address two questions: (1) what are the likely mtDNA groups of women taken from a given region (say, Birmingham); and (2) how genetically different are two mtDNAs taken from these groups?



(left) Due to the history and evolution of human populations, some mtDNA types -- denoted here by letters -- are historically more common in different world regions. (right) Our analysis of large-scale sequence data tells us how genetically different two mtDNAs from randomly-sampled women from different ancestral backgrounds are likely to be (circle size). The more different, the more likely the therapies involving that pair of women will experience difficulties.

To answer these, we retrieved (from the NCBI database) over 7000 human mtDNA sequences, as well as information about the mtDNA makeup of pre-industrial different regions around the world, and census information about the UK's, London's, and Birmingham's ethnic makeup. We used this information to estimate the mtDNA makeup of modern human populations -- which have become highly mixed through migration in recent times. Using these estimates, we then simulated thousands of Matilda-Wilma pairings in specific regions around the world (including the UK, London, and Birmingham). We recorded the genetic differences between these simulated pairs of mtDNAs to see how different we may expect women from different regions to be. The results have just appeared in Molecular Human Reproduction here; a similar, pre-peer-review version can be viewed for free here.

We found that the size of genetic differences likely to arise when sampling pairs women from modern populations was around 20-80 SNPs (single nucleotide polymorphisms -- specific molecular differences in mtDNA). This level of difference was enough to lead to substantial segregation bias in mouse models, suggesting that unprincipled choice of Wilmas from the general population could be problematic. These large differences are in large part due to modern population mixing, with substantial mixing of African and Asian mtDNA in modern UK cities contributing to the diversity. We showed that "haplotype matching" -- checking that Wilma is genetically similar to Matilda -- decreases these differences and so decreases the likelihood of problems with therapies. We also created a preliminary chart to help this process, showing which human haplotypes are genetically similar to others -- hopefully this will both help scientific understanding and therapeutic implementation in this field. Iain and Ellen

Thursday, 18 February 2016

ARTICLE: Who keeps the plans for our power stations?

Evolutionary Inference across Eukaryotes Identifies Specific Pressures Favoring Mitochondrial Gene Retention

IG Johnston, BP Williams
Cell Systems 2 (2), 101-111 (2016)

  • Why some genes are retained in mitochondria, where they are prone to disease-causing mutation, is a much-debated evolutionary question: we use new and generalisable maths and statistics to harness a large volume of sequence data and find the features of genes that predict the patterns of mitochondrial evolution that we observe.


Billions of years ago, a single-celled organism that would become our ancestor engulfed another smaller single-celled organism. The engulfed cell was probably intended to be lunch, but for reasons that remain mysterious (though recently explored here), it remained intact within our ancestor. It produced valuable chemicals that our ancestor could make use of, and was protected within the larger cell. This started a mutually beneficial relationship that evolved over billions of years to give rise to our situation today -- we are the descendants of the big cell, and our mitochondria are the descendants of the small, engulfed cell. 

As they were once independent organisms, mitochondria possess their own genomes (mitochondrial DNA, or mtDNA, which we've written about before). However, unlike the genomes of independent single-celled organisms like bacteria, mtDNA has only a handful of genes: why? Over evolutionary time, the majority of genes have either vanished from mtDNA or been transferred to the nucleus of the host cell. The reasons for transferring these genes to the nucleus are quite well understood; the nucleus is a safer environment for genes, less prone to mutation, and has several other evolutionary advantages.  But, given that transfer to the nucleus is possible, and genes in mtDNA are susceptible to mutation and damage (often giving rise to devastating diseases, which we study and try to prevent), why have mitochondria retained any genes at all? 

This question has been asked for decades, but until recently we lacked the data and the mathematical language to answer it quantitatively. Scientists energetically debate several different hypotheses: our approach attempts to let the data speak for itself without any preconceived ideas about which hypotheses are most likely. To this end, we built a mathematical model encompassing the evolutionary history of organisms with mitochondria, and a powerful statistical framework to amalgamate all the data that has been collected in recent years -- thousands of mitochondrial genomes from organisms from plants to protists (and humans) -- and harness it to compare the many disputed hypotheses addressing this question. 

Our mathematical approach allows us to "rewind the tape of evolution" and explore how mitochondrial genes have evolved. We're looking at Complex I -- an important protein complex involved in respiration -- over time, and watching the number of its subunits encoded in mitochondrial DNA (coloured black) decrease over evolutionary time, according to rules which we identify. The skyscrapers in the background are part of a graph describing how more mtDNA genes are lost over evolutionary history.

In a new paper in Cell Systems here (free here) we found several features that are most related to whether a gene is retained in mtDNA. Before discussing what they were, note that this picture -- several different features each with some influence -- explains and justifies the existing scientific debate. If hypothesis X and hypothesis Y both represent parts of the underlying "truth", then scientists advocating X alone and scientists advocating Y alone are neither completely wrong nor necessarily at odds -- everyone's partly right and the truth lies in the combination of the two arguments. 

The features that predict mtDNA gene retention are how central a gene's product is in its protein complex, the hydrophobicity of the protein the gene encodes, and the proportion of G's and C's in the gene's sequence. This suggests that genes are retained in mtDNA:  
(a) To allow local control of mitochondrial machinery (individual mitochondria can be controlled in response to cellular demands, rather than having to apply changes to the entire cellular population of mitochondria at once).  
(b) To prevent hydrophobic proteins ending up in the wrong place in the cell (if encoded by the far-away nucleus, these proteins may not be able to reach or enter the mitochondrion). 
(c and most speculatively) Because they are capable of withstanding the damaging environment of the mitochondria (GC-rich DNA and RNA is chemically more robust than GC-poor molecules). 

We found that the combination of the features we identified also predicted the success of experiments where scientists have attempted to mimic evolution and artificially transfer genes from the mitochondrion to the nucleus. Our results, as well as addressing a central mystery of evolutionary biology, thus also have the potential to inform synthetic biology approaches to tailor the genetics and bioenergetics of organisms. One final but important point is that the mathematical and statistical machinery we built for this project is highly generalisable and an efficient way of harnessing large sets of data about evolutionary and progressive processes -- we hope to use it to explore lots of other questions, including figuring out the pathways of disease progression and suggesting personalised medicine strategies in the clinic. Iain and Ben

Wednesday, 27 January 2016

ARTICLE: How evolution deals with mitochondrial mutants (and how we can take advantage)

Stochastic modelling, Bayesian inference, and new in vivo measurements elucidate the debated mtDNA bottleneck mechanism

  • Disease-causing mutant mtDNA is inherited through a complicated process: we use maths and statistics to shed light on this process and suggest possible therapeutic strategies to address disease inheritance and onset
Our mitochondrial DNA (mtDNA) provides instructions for building vital machinery in our cells. MtDNA is inherited from our mothers, but the process of inheritance -- which is important in predicting and dealing with genetic disease -- is poorly understood. This is because mitochondrial behaviour during development (the process through which a fertilised egg becomes an independent organism) is rather complex. If a mother's egg cell begins with a mixed population of mtDNA -- say with some type A and some type B -- we usually observe hard-to-predict mtDNA differences between cells in the daughter. So if the mother's egg cell starts off with 20% type A, egg cells in the daughter could range (for example) from 10%-30% of type A, with each different cell having a different proportion of A. This increase in variability, referred to as the mtDNA bottleneck, is important for the inheritance of disease. It allows cells with higher proportions of mutant mtDNA to be removed; but also means that some cells in the next generation may contain a dangerous amount of mutant mtDNA. Crucially, how this increase in variability comes about during development is debated. Does variability increase because of random partitioning of mtDNAs at cell divisions? Is it due to the decreased number of mtDNAs per cell, increasing the magnitude of genetic drift? Or does something occur during later development to induce the variability? Without knowing this in detail, it is hard to propose therapies or make predictions addressing the inheritance of disease.

We set out to answer this question with maths! Several studies have provided data on this process by measuring the statistics of mixed mtDNA populations during development in mice. The different studies provided different interpretations of these results, proposing several different mechanisms for the bottleneck. We built a mathematical framework that was capable of modelling all the different mechanisms that had been proposed. We then used a statistical approach called approximate Bayesian computation to see which mechanism was most supported by the existing data. We identified a model where a combination of copy number reduction and random mtDNA duplications and deletions is responsible for the bottleneck. Exactly how much variability is due to each of these effects is flexible -- going some way towards explaining the existing debate in the literature.  We were also able to solve the equations describing the most likely model analytically. These solutions allow us to explore the behaviour of the bottleneck in detail, and we use this ability to propose several therapeutic approaches to increase the "power" of the bottleneck, and to increase the accuracy of sampling in IVF approaches.

A "bottleneck" acts to increase mtDNA variability between generations. But how is this bottleneck manifest? Our approach suggests that a combination of copy number reduction (pictured as a "true" copy number bottleneck), and later random turnover of mtDNA (pictured as replication and degradation), is responsible.

Our excellent experimental collaborators, lead by Joerg Burgstaller, then tested our theory by taking mtDNA measurements from a model mouse that differed from those used previously and which, could in principle have shown different behaviour. The behaviour they observed agreed very well with the predictions of our theory, providing encouraging validation that we have identified a likely mechanism for the bottleneck. New measurements also showed, interestingly, that the behaviour of the bottleneck looks similar in genetically diverse systems, providing evidence for its generality. You can read about this in the free (open-access) journal eLife here. Iain and Nick [blog article also here]

ARTICLE: Evolutionary competition within our cells: the maths of mitochondrial DNA

mtDNA Segregation in Heteroplasmic Tissues Is Common In Vivo and Modulated by Haplotype Differences and Developmental Stage


  • MtDNA mixtures in cells arise through mutation and gene therapies: we show that different types of mtDNA usually proliferate at different rates, which suggests ways that therapies could be made more efficient.
Women may carry mutated copies of mitochondrial DNA (mtDNA) -- a molecule that describes how to build important cellular machinery relating to cellular energy supply. If this mutant mtDNA is passed on to that woman's child, the child may develop a mitochondrial disease, which are often degenerative, fatal, and incurable.

Joerg created mice that contained two types of mtDNA -- here illustrated as blue (lab mouse mtDNA) and yellow (mtDNA from a mouse from a wild population). We used several different wild mice from across Europe to represent the mtDNA diversity one may find in a human population. We found that throughout a mouse's lifetime, one mtDNA type often outcompetes another (here, yellow beats blue), with different patterns across different tissues.

Amazing new therapies potentially allow a carrier mother A and a father B to use another woman C's egg cells to conceive a baby without much of mother A's mtDNA being present. The approach involves taking nuclear DNA content from A and B (so that most of the child's features are inherited from the true mother and father), and placing it into C's egg cells, which contain a background of healthy mtDNA. You can read about, what are misleadingly called, three-parent babies here.

Something that is less discussed is that, in this process, a small amount of A's mutant mtDNA can be "carried over" into C's cell. If this small amount remains small through the child's life, there is no danger of disease, as the larger amount of healthy C mtDNA will allow the child's cell to function normally. We can think of the resulting situation as a competition between A and C -- if A and C are evenly matched, the small amount of A will remain small; if C beats A, the small amount of A will disappear with time; and if A beats C, the small amount of A will increase and may eventually come to dominate over C.

Until recently it has been fair to assume that A and C are always about evenly matched (unless something is drastically different between A or C). However, evidence for this idea was based on model organisms in laboratories, which do not have the same amount of genetic diversity as found in human populations. Our collaborator Joerg addressed this by capturing wild mice from across central Europe, selecting a set that showed a comparable degree of genetic diversity to that expected in a human population. He used these, with our modelling and mathematical analysis, to show that pronounced differences between A and C often exist, and are more likely in more diverse populations. The possibility that A beats C, and mutant mtDNA comes to dominate the child's cells, therefore cannot be immediately discounted in a diverse population. We propose "haplotype matching" -- ensuring that A and C are as similar as possible -- to ameliorate this potential risk. It's open as to whether one can generalize from observations in mice to people and it's also open as to whether our conclusions, which used lab-mice as parent A (which are not entirely typical creatures) of necessity generalize to other non-lab mouse types.

Our mathematical approach also allowed us to explore, in detail, the dynamics by which this competition within cells occurs. We were able to use our data rather effectively by having a statistical model that allowed us to reason jointly about a range of data sets. We found that the degree to which one population of mtDNA beat the other depended on how genetically different they were.  We found that different tissues were like different environments: some favouring C over A and some vice-versa. This is perhaps surprising to some as this evolution in the proportions of different genetic species is not something we imagine occurring inside us, during our lives, and as something that might differ between our organs. We found several different regimes, where the strength of competition changes with time and as the organism develops: when our cells are multiplying faster they show a more marked preference for one of the species. We've shown our results to the UK HFEA in its ongoing assessment of these therapies, and you can read, for free, about our work in the journal Cell Reports here. Iain, Joerg, Nick [blog article also here].

ARTICLE: Polyominoes: mapping genotypes to phenotypes

A tractable genotype–phenotype map modelling the self-assembly of protein quaternary structure


  • Proteins in our cells have intricate structures built by genetic instructions, and these structures are vital for life: we produce a computational model to explore the relationship between genetic instructions and structure, and how evolution and mutation may change proteins.
Biological evolution sculpts the natural world and relies on the conversion of genetic information (stored as sequences, usually of DNA, called genotypes) into functional physical forms (called phenotypes). The complicated nature of this conversion, which is called a genotype-phenotype (or GP) map, makes the theoretical study of evolution very difficult. It is hard to say how a population of individuals may evolve without understanding the underlying GP map.

This is due to the two fundamental forces of evolution -- mutations and natural selection -- acting on different aspects of an organism. Mutations occur to genotypes (G), while natural selection, the ultimate adjudicator of the fate of mutations in the population, acts on the phenotype (P). Without understanding the link between these two -- the GP map -- we can't easily say, for example, how many mutations we expect important proteins within a virus strain to undergo with time, and thus how quickly the virus will evolve to be unrecognised by our immune systems.

Simple models for the mapping of genotype to phenotype have helped answer important questions for some model biological systems, such as RNA molecules and a coarse-grained model of protein folding. One important class of biological structure which has not yet been modelled in this way are protein complexes: structures formed through proteins binding together, fulfilling vital biological functions in living organisms. In this work, we introduce the "polyomino" model, based on the self-assembly of interacting square tiles to form polyomino structures. The square tiles that make up a polyomino are assigned different "sticky patches", modelling the interactions between different proteins that form a complex. A huge range of structures can be formed by varying the details of these patches, mimicking the range of protein complexes that exist in biology (though there are some obvious differences in the shapes of structures that can be formed).

Our simple model explores the interactions between protein subunits, and how these interactions shape a surface that evolution explores. (top) Sickle-cell anemia involves a mutation that changes the way proteins interact, making normally independent units form a dangerous extended structure. (bottom) Our polyomino model models this effect. The resultant dramatic effects on structure, fitness, and evolution can then be explored.

Despite its abstraction we show that the polyomino model displays several important features which make it a potentially useful model for the GP map underlying protein complex evolution. On top of this, we demonstrate that our model possesses similar properties to RNA and protein folding models, interestingly suggesting that universal features may be present in biological GP maps and that the "landscapes" upon which evolution searches may thus have general properties in common. You can find the paper free here and you can play with polyominoes here! Iain [blog article also here]

ARTICLE: Inferring the evolutionary history of photosynthesis : C 4 yourself


Phenotypic landscape inference reveals multiple evolutionary paths to C4 photosynthesis

  • Some plants have evolved efficient photosynthesis, but important crops including rice have not: we use maths and statistics in conjunction with biological data to understand this evolution, with a view to repeating it artificially in crops to increase food production

Biological evolution is a complex, stochastic process which dictates fundamental properties of life. Our understanding of evolutionary history is severely limited by the sparsity of the fossil record: we only have a handful of fossilised snapshots to infer how evolution may have progressed throughout the history of life. Many physicists and mathematicians have attempted theoretical treatments of the process of evolution, using varying degrees of abstraction, in order to provide a more solid quantitative foundation with which to study this complex and important phenomenon, but the predictive power of these theoretical models, and their ability to answer specific biological questions, is often questioned.

We recently focussed on one remarkable product of evolution in plants: so-called "C4 photosynthesis". C4 consists of a complex set of changes to the genetic and physiological features which have evolved in some plants and act to increase the efficiency of photosynthesis. This complex set of changes has evolved over 60 times convergently: that is, plants from many different lineages independently "discover" C4 photosynthesis through evolution. We were interested in the evolutionary history of how these discoveries occurred -- both motivated by fundamental biology and the possibility of "learning from evolution" and using information about the evolution of C4 to design more efficient crop plants.


C3 and C4 plants differ in several physical and genetic ways (leaves and cells either side). We picture evolution as progressing along paths over a hypercube connecting these states (grey lines) -- some paths will give rise to intermediate species matching those we really observe (red and blue points). We can calculate how likely each path is and thus reconstruct evolutionary history.
 
To this end, we modelled the evolution of C4 as a pathway through a space containing many different possible plant features. The pathway starts at C3 -- the precursor to C4 -- and progressively takes steps in different directions, acquiring one-by-one the features that sum up to C4 photosynthesis. Using a survey of plant properties from across the wide scientific literature, we identified which intermediate states these pathways were likely to pass through, given observed properties of plants that currently possess some, but not all, C4 features. We were then able to use a new inference technique to predict the ordering in which these likely pathways traverse the evolutionary space. We showed that this approach worked by both successfully inferring the known evolutionary steps in synthetic datasets and correctly predicting previously unknown properties of several plants, which we verified experimentally. Our (open access) paper is here and there's a less technical summary and commentary here. Our approach showed that C4 photosynthesis can evolve through a range of distinct evolutionary pathways, providing a potential explanation for its striking convergence. Several of these different pathways were made explicitly visible when we examined the inferred evolutionary histories of different plant lineages -- different families are likely to have converged on C4 through different evolutionary routes. Furthermore, the most likely initial steps towards C4 photosynthesis are surprisingly not directly related to photosynthesis, being solutions to different biological challenges, but also providing evolutionary "foundations" upon which the machinery of C4 can evolve further. We hope that the recipes for C4 photosynthesis that we have inferred find use in efficient crop design, and anticipate our inference procedure being of use in the study of other specific biological questions regarding evolutionary histories. Iain [blog article also here]

ARTICLE: Walking on evolutionary landscapes

Epistasis can lead to fragmented neutral spaces and contingency in evolution

  • Mutations can cause disease and diversity by changing structures in biology: we use a computational model to understand what changes are possible and how evolutionary future is constrained by the current state of organisms
Evolution can, in an abstract sense, be pictured as a journey through a genetic "space". A set of co-ordinates (like latitude and longitude, but more detailed) in this space corresponds to an organism's "genotype" -- the ordered set of As, Cs, Gs, and Ts found in its DNA. Each genotype encodes information about physical features of the organism -- its "phenotype". Mutations and other genetic changes cause steps from one point to another in genome space, and some of these steps will change the phenotype of an organism. As an artificially simple example, imagine a case where the genotypes AA and AC may give an organism red feet, but AG and AT give it blue feet. The offspring of a red-footed organism with genotype AA may pick up an A->C mutation in the second position (AA->AC) and keep their parent's red feet; or they may pick up an A->T mutation (AA->AT) and have a new blue-footed phenotype.

The structure of this evolutionary space -- the pattern of phenotypes encoded by connected genotypes -- clearly affects how mutations can change the form of organisms, and is thus central to our understanding of evolution. Mutational changes cause genetic diseases, allow bacteria and viruses to adapt to our immune response, and generate the beautiful biodiversity in the world around us. We aim to learn more about these systems, how they change with time, and their evolutionary limitations, by studying evolutionary spaces in a computer.

A schematic genome space. Each point on the grid is a genotype, which encodes a phenotype (red squares, blue circles, green diamonds, etc). Evolution can step between adjacent points through mutations -- a mutation may move an organism to its left neighbour, for example, or upwards by one point. Some mutations -- for example, a rightwards step from the top left corner -- keep the phenotype (green diamond) intact. Some (a downwards step from the top left corner) change the phenotype (green to blue). Not all genotypes encoding the same phenotype are connected, and different clusters have different evolutionary potential. For example, a gold triangle encoded by the cluster in the top right can only stay gold or become blue; a gold triangle encoded by the cluster on the left can stay gold or become blue, grey, or red.

We chose to look at the evolutionary space of RNA -- a class of biological molecule, examples of which play many vital roles in our cells. RNA, like DNA, consists of an ordered series of chemical groups denoted by letters, and RNA molecules fold into particular structures governed by this sequence of letters. These structures are central to the function of some RNAs, and the structure can be predicted by a computer program from the sequence of letters. So we have a model system where the phenotype (structure) corresponding to a genotype (letters) can easily be computed.

We explored the full genetic space of RNA molecules that consist of 15 letters (meaning that our computer had to fold over a billion structures!). In our survey in Proceedings of the Royal Society B here (free here) we found a large skew in the numbers of genotypes that encode a phenotype -- some structures are encoded by many different sets of letters and occupy a vast amount of the genetic space, and some are encoded only by very few genotypes. As mutations are random, we may expect to see these more frequent structures more commonly in the natural world (if all structures are affected equally by selective pressures). We also found that sets of genomes encoding the same phenotype are often disconnected in genome space. To pursue our simple example above, imagine AA and AC give red feet, AG and AT blue feet, TT and TG red feet and GT and GG green feet. The AA/AC red genomes aren't connected by single mutations to the TT/TG red genomes -- they form separate "clusters" in genospace. Furthermore, while a redfoot with genotype TT or TG can mutate to become either blue-footed (T->A in the first position) or green-footed (T->G in the first position), a redfoot with genotype AA or AC can only mutate to become blue-footed and cannot access the greenfoot genotypes without changing to something else first. One can see that this disconnected nature of genospace makes evolution contingent on genotype: redfoots with different genotypes can change in different ways. Understanding how this contingency appears in different systems will help us describe and even predict the outcomes of evolutionary processes. Iain

Tuesday, 26 January 2016

ARTICLE: Pretty polyominoes

Evolutionary dynamics in a simple model of self-assembly

  • The evolution of self-assembling structures in biology is hard to study: we build and explore a model that contains key features (a genome encoding interactions between physical subunits) and how evolution "learns" self-assembly rules
Lots of vital structures in biology must build themselves in the chaotic, frantic environment of our cells. Proteins -- machines in our cells that perform all sorts of tasks, from building cell scaffolding to assisting chemical reactions and moving electric charges -- are no exception. Many proteins self-assemble in cells, with different subunits coming together and sticking to each other to form a functioning product. The precision with which this assembly occurs "automatically" has been metaphorically described as a hurricane blowing through a pile of Lego and a fully-formed Lego train forming.

Proteins are complicated structures that have evolved over billions of years. To understand how evolution has "learned" the interactions between subunits that proteins need to successfully self-assemble, we need to work with a model system that includes the scientifically interesting details but is simple enough to investigate on paper or with a computer.

In a paper (here in Physical Review E; free here), we work with "polyominoes" as a model for evolving self-assembling systems. Polyominoes are shapes made up of connected square tiles (dominoes are polyominoes containing two tiles; the "tetromino" pieces in Tetris -- from which the game gets its name -- are polyominoes containing four tiles). In our model, these tiles have sticky edges, with some edges sticking to some others -- a list of numbers describes who sticks to whom. If we put tiles in a bag and shake them, some edges will stick, and a polyomino will form. We thus have the essential features of a random cell (the bag) and interactions between subunits (sticky edges) which are encoded by a genome (the list of numbers).




(left) The polyomino model. A "genome" (list of numbers) describes a set of square tiles with patchy edges. Some patches stick to each other -- in this case, 1 sticks to 2, 3 sticks to 4, and so on (0 sticks to nothing). Then, if the tiles are allowed to mix and meet, these sticky patches meet and form a larger structure. (right) The range of structures (axes) that can be built using two tile types, and the ways that evolution, acting on the model genomes, can change between output structures through single mutations (pink -- no link; darker -- more possible transitions).

We then simulated evolution in a computer, with random mutations changing the "genome" and natural selection acting on the resulting polyominoes. We show how different mutation rates, population sizes, and reproductive strategies affect the evolution of model biological structures. We also show that this evolutionary model can explain the symmetries of protein structures observed in biology. We explore the surprisingly rich set of structures that the model can produce, and the different modes of evolution that can lead from simple to complex forms -- helping us understand how important structures have evolved and may go wrong through mutation. There's more work on this here and an associated interactive simulation tool here! Iain

ARTICLE: Evolving social networks of genes

The effect of scale-free topology on the robustness and evolvability of genetic regulatory networks

  • How life both diversifies and maintains its function in the face of random mutation is an open question: we model the networks describing interactions between genes to explore this evolutionary tradeoff
A big question in evolutionary biology centres around an apparent paradox. To prevent mutations (which inevitably occur throughout life) from having damaging or fatal effects, organisms must be "robust" -- they must retain their biological functions even if some mutations occur. But to be able to adapt to changing environments and situations as generations pass, they must also be "evolvable" -- mutations must be able to change aspects of their biological functionality. How can we have a situation where mutations both have no functional effect and have the ability to change functionality?

To explore this question, we consider a class of key players in biological functionality: gene regulatory networks (GRNs), which describe how genes regulate each others' production in our cells. The products of one gene may lead to increased or decreased expression of another gene's products; this regulation is used by the cell to control its contents in response to signalling and sensory inputs.

In this paper, in the Journal of Theoretical Biology here (free here), we investigate the effects of mutations on the functions of model GRNs. Specifically, we calculate the different types of behaviour that a model GRN can show -- some of these involve a fixed state where every gene is either on or off, and some involve a cycling process where genes periodically switch on then off in a fixed pattern. We modelled mutations by randomly removing parts of the network, and explored how these mutations changed the GRN behaviour. If a random mutation changed the behaviour of the GRN (for example, changing the switching patterns of a gene), it is more evolvable; if the behaviour stays the same after a mutation, it is more robust.


(top) Several model GRNs: dots are genes, arrows describe one gene enhancing the production of another; flat ends describe one gene decreasing the production of another. (bottom) The gene states that each network can experience. Each point is a different on/off pattern of genes; the loops in the centre of the "flowers" are cycles of gene states that all connected states eventually collapse down to over time.

We found that the structure of a GRN dramatically affects the influence of mutations. We investigated so-called "Erdos-Renyi" (ER) structures, where nodes are connected randomly, and "scale-free" (SF) structures, where some "hub" nodes are connected to lots of others, but many more nodes connect only to very few neighbours. Social networks are often SF, with a small number of highly connected people and a large number of less connected one. SF networks were both more evolvable and more robust in the face of random mutations than ER ones: fewer mutations changed their behaviour, but those mutations that did have an effect gave rise to a more diverse set of new behaviours. We also found that SF networks were more robust to changes in environment (random changes to individual gene states). In biology we often observe that GRNs are scale-free: this work suggests the evolutionary advantages enjoyed by this class of networks, and thus a possible explanation for their appearance in biology. Iain


Sunday, 24 January 2016

ARTICLE: DNA in computers

The self-assembly of DNA Holliday junctions studied with a minimal model

  • Physically rich DNA structures are vital in biology and genetics, and valuable in nanotechnology: we show that they can be studied with coarse-grained computer models, which are simple enough to simulate but rich enough to match real-world behaviour
DNA is a fascinating and flexible model. In our cells, it forms a famous "double helix" structure, but lots of other structures too, including "Holliday junctions", four-armed crosses that occur when DNA molecules meet and exchange genetic information. DNA is also used in nanotechnology, where our ability to design DNA molecules which interact in designed ways is harnessed to produce tiny molecular structures and machines. Holliday junctions also play important roles in this DNA nanotechnology, forming the corners of rigid structures.

To understand the physics of how DNA behaves in both these biological and nanotechnological contexts, we explored whether a simple model of DNA simulated in a computer can describe and predict its behaviour in the real world. DNA is made of many atoms and interacts in a complicated way with its environment: we aimed to reduced these complications as far as possible while retaining the necessary information to describe the physics of interest.

In a paper in the Journal of Chemical Physics here (free here), we build a model of DNA using common pieces of "kit" from a physicists' toolbox: model sticky particles linked in a chain, with interactions limiting how much the chain can be bent and twisted. The model is simple enough to simulate easily on a computer, but correctly forms duplexes and Holliday junctions analagous to those we see in the real world. It's a demonstration that so-called "coarse-grained" models (as opposed to modelling in fine detail) can be of use in simplifying and understanding complicated structures and physical systems.

Model DNA molecules forming four-armed Holliday junctions in a computer simulation.

This philosophy has since been developed and expanded hugely, leading to the highly influential oxDNA project, which has been used to explore how DNA nanomachines move and function and has led to a wide range of insights in chemical physics and nanotechnology. Iain