Showing posts with label probability and stochastic processes. Show all posts
Showing posts with label probability and stochastic processes. Show all posts

Wednesday, 23 February 2022

ARTICLE: Stochastic fantasy combat!

Optimal strategies in the Fighting Fantasy gaming system: influencing stochastic dynamics by gambling with limited resource
Iain G Johnston
European Journal of Operational Research DOI 10.1016/j.ejor.2022.01.039 (2022)

Here's a more unusual one. Fighting Fantasy gamebooks, hugely popular in the 80s and resurgent now, are "games in a book". The reader/player chooses their path through the book's many sections, overcoming challenges, fighting monsters, and hopefully succeeding in their quest.

The combat system is quite interesting. The player and their opponent both have "stamina" -- like the health bar in a video game. The player rolls dice to determine the strength of one of their attacks, and rolls again for the opponent. Whoever has the higher strength inflicts some stamina loss on the other. The combat proceeds through rounds like this until someone's stamina reaches zero.

Phrased like this, the player has no agency and the system is quite easy to solve (ie, work out the probability of winning a given fight). But there's another factor. The player (not the opponent) also has some "luck", describing how lucky they are. Testing luck involves rolling two dice: if the sum is less than or equal to the player's luck score, they are luck, otherwise they're unlucky. If they win an attack round, they can choose to test their luck, and if luck they deal more damage to their opponent. If they lose an attack round, they can also test their luck, and will take less damage if lucky. If they're unlucky, the outcome is negative: they do less, and take more, damage than if they'd not tested.

A bit more complicated, but still possible to solve. But here's the rub. Every time you test your luck, your luck score decreases -- whether you're lucky or unlucky. So as you test your luck more and more, you are less and less likely to get a positive outcome. The question is -- when is it a good idea to use a luck test in combat?

To answer this we used an approach called dynamic programming. We first considered all the ways combat can end -- with someone having no stamina left. We next considered every state of the combat that could lead to one of these end states, and calculated the probability of each possible outcome in the case where the player chose to test their luck and the case where they didn't. We then considered all states of combat that led to these states, and so on, multiplying probabilities as we went to calculate the overall probability of victory from any state given any luck strategy.

We found that judicious use of luck can dramatically increase victory probability in some cases, particularly when player and opponent statistics are unbalanced. There are some general rules -- for example, no matter how low your luck, you should always test if you are otherwise about to die. We used some tools from statistics to distil the complex set of detailed optimal strategies into more general principles to follow. We also connect to more real-world questions, like cheating and espionage, where a "lucky" outcome can be beneficial -- but an unlucky one can be disastrous, and the more you test your luck the more likely you are to be unlucky!

ARTICLE: Corals to crops -- how life protects the plans for its cellular power stations

Avoiding organelle mutational meltdown across eukaryotes with or without a germline bottleneck
David M Edwards, Ellen C Røyrvik, Joanna M Chustecki, Konstantinos Giannakis, Robert C Glastad, Arunas L Radzvilavicius, Iain G Johnston
PLoS Biology 19 e3001153 (2021)

(this text is from a press release about the article)

An international team of researchers led by the University of Bergen has uncovered how organisms from crops to corals may avoid deadly DNA damage during evolution.

Our cells, and those of animals, plants and fungi, contain compartments that produce chemical fuel. These compartments contain their own DNA, which stores instructions for important cellular machinery. But this so-called oDNA (organelle DNA) can become mutated, corrupting the instructions and preventing cells making enough energy.

In humans and some other animals, a process called the “bottleneck” allows some offspring to inherit less mutated oDNA. This process needs mothers’ egg cells to develop early, like in humans, where a human girl is born with all her egg cells already formed. But other organisms, from plants to fungi, don’t develop these cells early – their flexible body plans mean that eggs are not “set aside” early in development.<\p>

“We wanted to know how these organisms might avoid inheriting mutations without a human-like bottleneck,” said Ellen Røyrvik, a geneticist on the research team, based at UiB.

The scientists used mathematical modelling to show that a process called gene conversion – the controlled overwriting of DNA – could in theory allow some offspring to inherit less mutant oDNA without requiring a bottleneck. Using genome data, they found machinery controlling this process in plants and fungi, but also in soft corals, sponges, and algae – all organisms without fixed body plans. They also found that this machinery was most active in the parts of plants that will end up producing the seeds of the next generation, suggesting that it is indeed used to allow some offspring to inherit fewer mutations.

Organisms without fixed body plans (including octocorals, sea pens, sponges, plants, and fungi) and with fixed body plans (including humans and many animals) may use different strategies to avoid the buildup of damage in their cellular "power stations." CREDIT: Gemma Lofthouse

“Taken together, it looks like organisms without a fixed body plan – plants, fungi, corals, sponges, algae – may have adopted gene conversion to deal with oDNA mutations,” said Iain Johnston, an associate professor in the Mathematics Institute at UiB, who led the research. “Humans and other animals can develop egg cells early and use a bottleneck; other organisms can use gene conversion instead.”

Going forward, the team plans to explore how this overwriting of oDNA causes other issues in the organisms that use it – including crop plants, where it can cause sterility. They are also exploring the broader question of why these compartments contain oDNA at all, given the risk of mutational damage.

Friday, 23 April 2021

ARTICLE: How does tool use evolve in animals?

Data-Driven Inference Reveals Distinct and Conserved Dynamic Pathways of Tool Use Emergence across Animal Taxa, iScience 23 101245 (2020)

There are some wonderful examples of animals using tools. Octopuses block oyster shells open with coral; boxer crabs wave captive anemones for defense and food capture; captive dolphins use feather to wipe clean their aquarium windows. A while ago, we saw this excellent infographic in National Geographic, and got interested in these data. Do different families of animals learn to use tools in completely different ways? Or are there some general ("universal") principles behind how animals learn to use tools?

There is, of course, a big and fascinating literature on tool use, but we found rather few studies attempting a quantitative comparative analysis across bilaterian animals. To address other evolutionary questions, we've developed HyperTraPS (hypercubic transition path sampling), a statistical approach for learning the "pathways" of evolutionary processes. That is, which events occur before and after which other events in an evolving system? Does feature A always evolve before feature B? We used HyperTraPS to ask about the orderings with which different types of tool use appeared in animals. For example, do animals always learn to "poke" before they learn to "dig"? Do all animals learn tool use in the same order, first A then B then C..., or does it vary across species?

We found some answers that we think are quite interesting. There seem to be some quite deep similarities across animal species in how tool use evolves. Types of tool use like "affixing" and "throwing" are almost universally acquired early; types like "cutting" and "symbolising" are acquired late and rarely, only by primates. The environment and animal family influences the structure of these pathways: aquatic organisms seem to discover "waving" tool use relatively early, for example, and primates discover tools that "block" relatively late.

(A) The inferred pathways of tool use emergence across animals. The size of a blob gives the probability that that mode of tool use (on the horizontal axis) is acquired at that stage (on the vertical axis) of a species' discovery of tool use types. (B) Sample evolutionary pathways of tool use, with individual animal lineages illustrated at the positions corresponding to the modes of tool use they have discovered.

Of course, there's a lot of uncertainty about any analysis like this. Are we talking about wild or captured animals? What if we just haven't observed some types of tool use? We attempted to address several such questions with our analysis and showed that our overall results were quite robust with respect to these uncertainties. HyperTraPS fully describes the uncertainty in its outcomes, helping interpretability. We hope that our results help at least to suggest some possible principles and points for further investigation in this fascinating topic. You can read more in iScience here.

ARTICLE: What makes mitochondria selfish, and when do selfish ones win?

MtDNA sequence features associated with ‘selfish genomes’ predict tissue-specific segregation and reversion, Nucleic Acids Research 48 8290 (202)

Mitochondria, the power stations of the cell, are in some senses like people in a company. The company needs several people to contribute if it is to survive. Some people may work hard and contribute lots to the company. Others may selfishly slack off and rely on others doing the work.

The cell needs mitochondria to produce ATP, the chemical that powers many important processes. But there is evidence that some mitochondria are more selfish, and some less so, than others. Unselfish mitochondria produce machinery which helps produce ATP. Selfish mitochondria prefer to replicate, copying their DNA and contributing less to the cell. Interestingly, at the molecular level, there is something like a "switch": a mitochondrion either takes steps to produce useful machinery, or takes steps that will help it replicate.

We were interested in why different mitochondria choose different positions of this switch, and what is a good "strategy" for mitochondria under different conditions. We built a simple model of this behaviour to understand it. Unsurprisingly, we found that selfish mitochondria -- favouring replication, and contributing less to the cell -- profilerate over unselfish ones in cells where there's little pressure to co-operate. As they replicate more, selfish mitochondria eventually come to dominate such cells. Where there is cellular pressure, however, unselfish mitochondria may win out. This is because cells full of selfish mitochondria won't perform adequately, and the whole cell and all its mitochondria will die -- leaving those cells with more unselfish mitochondria remaining.

What do we mean by "cellular pressure"? Well, if some type of cells never die, clearly the latter event can't happen, and we'd expect selfish mitochondria to win. If cells die regularly, perhaps there's more capacity to select those filled with unselfish mitochondria. We looked at different tissues where cells die with different rates, in mice where cells had two different types of mitochondrial DNA (mtDNA). We found a consistent pattern where one type of mtDNA proliferated in slow-dying cells and the other proliferated in fast-dying cells. Why different mtDNA types "win" in different tissues is a big question (which we've looked at before!), and it looks like this might help explain some of these differences.


(A) Sequence features may make different mtDNA types more "selfish" (favouring replication) or "unselfish" (favouring the production of useful machinery). (B) Our theory shows how, depending on cellular pressures, one or the other strategy can be favoured, leading to selection for one or the other mtDNA type.

We also asked what it is about a particular mtDNA sequence that might make it more or less selfish. Based on how mtDNA produces useful machinery, and how it replicates, we hypothesised that some features in the so-called "control region" of mtDNA may influence selfishness. Using sequence information, we found that these features tied quite neatly in with the observations in these mouse models, and also in (more limited) observations from human cells. While certainly not resolved, this picture suggests a link between sequence features of mtDNA, cellular selfishness, and proliferation differences across different tissues. You can read more (for free) in Nucleic Acids Research here.

Wednesday, 8 January 2020

ARTICLE: Learning pathways of disease progression

HyperTraPS: Inferring probabilistic patterns of trait acquisition in evolutionary and disease progression pathways
Sam F Greenbury, Mauricio Barahona, Iain G Johnston
Cell Systems (2019)

Many diseases that take a substantial human toll can be viewed as “progressive”. That is, a patient starts out healthy, then disease-related problems and/or symptoms develop over time. For example, a given case of cancer may begin with a patient acquiring a particular mutation, then other mutations building up in their genome over time.

How the same disease progresses in different patients often varies widely. Understanding this variability is important for precision medicine, where detailed knowledge of individual patients is used to design the best targeted treatments. However, learning the varied pathways of diseases and using them to predict future outcomes is challenging. Human researchers usually cannot hope to remember or analyse enough examples of patient data to provide the most reliable picture.

We previously developed an algorithm called HyperTraPS (hypercubic transition path sampling) to explore how biological systems evolve over time. We reasoned that HyperTraPS could also be used to learn the pathways of disease progression. In a new study in Cell Systems (free preprint available here) we used HyperTraPS to analyse biomedical data from many patients – hundreds, or thousands of individuals – to build a ‘road map’ of the different pathways that a disease takes over time.

Picture a river that branches out into a wide delta. Patients start out healthy – upstream in the river – and different patients go down different branches as the disease progresses and they acquire more symptoms. HyperTraPS learns the structure of the river delta from data, and predicts which river branches are more or less likely – and, importantly, where you'll end up if you're currently at a particular point.

By learning these branching patterns of disease progression, HyperTraPS has helped provide a refined risk assessment for malaria, based on data from thousands of Gambian children – as we’ve written about before. The approach also revealed diverse pathways of ovarian cancer progression, where the first mutation to occur appears to play a large role in determining subsequent mutations.

The "waterfall" in the foreground shows paths from one stage of a disease to the next, learnt by HyperTraPS using data from a high number of patients. Each dot of the illustration represents different stages of disease, for example a specific set of symptoms or a given set of mutations. The thickness of the lines indicate the probability of moving from one specific stage of disease to the next.

HyperTraPS is very generalisable and can be used to learn pathways by which mutations, symptoms, or other features develop over time from an initial state. We further used this generalisability to understand a biomedically important example of evolution – specifically, how tuberculosis evolves to become resistant to antibiotics.

Tuberculosis acquires resistance through mutations, and HyperTraPS has revealed the patterns of these mutations in TB bacteria reported from a group of 1000 Russian patients. These patterns help predict which mutation a bacterium will acquire next, and hence which drugs may be more effective for a given case. We’re following up with other applications of HyperTraPS, to learn about other progressive diseases, ageing, and evolution, and even to analyse how students complete tasks in online courses.

ARTICLES: Evolving cellular populations of mtDNA

Evolving mtDNA populations within cells
Iain G Johnston, Joerg P Burgstaller
Biochemical Society Transactions 47 1367 (2019)
and
Varied mechanisms and models for the varying mitochondrial bottleneck
Iain G Johnston
Frontiers in Cell and Developmental Biology 7 294 (2019)

We've recently written two review papers looking at the dynamics of mitochondrial DNA (mtDNA) in cells. As we've written about before, cells contain populations of hundreds or thousands of mtDNA molecules. These molecules replicate and degrade, so that over time, cellular populations of mtDNA change and evolve. The amount of disease-causing mutations, and the number and structure of mtDNA molecules, may all change as organisms develop and age, with different consequences.

The first article, in Biochemical Society Transactions, takes a broad look at how cellular mtDNA populations change over time, considering a range of organisms from humans and other animals to plants and fungi. We look at the different processes that change mtDNA populations, which include replication and degradation but may also include recombination (particularly in plants) and cell-to-cell exchange of mitochondria. The review particularly highlights the importance of understanding cell-to-cell variability in mtDNA populations -- as it only takes a few cells with lots of mutant mtDNA to cause disease, it's important to understand the statistics of mtDNA populations across cells. We review experiments and theory aiming to do so, including our recent work showing that cell-to-cell variability of mtDNA mutant load increases over time in a wide variety of circumstances.

The second article, in Frontiers in Cell and Developmental Biology, focuses on the so-called "mtDNA bottleneck", a process that shapes mtDNA populations in early mammalian development, and helps prevent the inheritance of mutant mtDNA. Specifically, mtDNA undergoes a "genetic bottleneck" between generations, meaning that mothers' egg cells, and offspring, often have dramatically different mtDNA populations. The review emphasises that this "genetic bottleneck" is an effective quantity, not a directly measurable observation, that arises from several physical processes, including but not limited to a "physical bottleneck" or mtDNA depletion during development. Different ways of modelling, analysing, and explaining the "genetic bottleneck" are reviewed, from human populations to mouse egg cells and Adélie penguins. We invest some time in trying to explain the different assumptions, symbols, and methods that researchers have used to quantify the bottleneck over the years. Again, the importance of understanding cell-to-cell variance in mtDNA populations is a core theme.

A. Different processes shaping mixed mtDNA populations inside cells. B. The "genetic bottleneck", increasing mtDNA variance between egg cells and offspring.  

Like all reviews, these articles don't have new results, but attempt to summarise existing knowledge and thinking on these topics. We hope that both papers provide some interesting insights, references, and (in the case of the bottleneck paper) visualisations that may help understand these often confusing topics.



Monday, 15 July 2019

ARTICLE: Phenotypes and progression pathways in severe malaria

Precision identification of high-risk phenotypes and progression pathways in severe malaria without requiring longitudinal data
Iain G Johnston, Till Hoffmann, Sam F Greenbury, Ornella Cominetti, Muminatou Jallow, Dominic Kwiatkowski, Mauricio Barahona, Nick S Jones, Climent Casals-Pascual
npj Digital Medicine 2 63 (2019)





We recently published an article here in npj Digital Medicine using maths (including HyperTraPS) to learn more about severe malaria, a disease that kills over 400 000 people (mainly African children) a year. Severe malaria is challenging in the clinic because its symptoms and progress vary a lot from patient to patient. Our approach helps learn about this variability and identify high-risk patients and pathways. You can read our blog article about the paper on the npj Digital Medicine community blog here!



ARTICLE: The cell's power station policies


Energetic costs of cellular and therapeutic control of stochastic mitochondrial DNA populations
Hanne Hoitzing, Payam A Gammage, Lindsey van Haute, Michal Minczuk, Iain G Johnston, Nick S Jones


(Hanne's also written a post about this paper, you can read it here)

Our cells are filled with populations of mitochondrial DNA (mtDNA) molecules, which encode vital cellular machinery that supports our energy requirements. The cell invests energy in maintaining its mtDNA population, like us using electricity-powered tools to help maintain our power stations. Our cellular power stations can vary in quality (for example, mutations can damage mtDNA), and are subject to random influences. How should the cell best invest energy in controlling and maintaining its power stations? And can we use this answer to design better therapies to address damaged mtDNA?

In a new paper here in PLoS Computational Biology, we attempt to answer this question using mathematical modelling, linking with genetic experiments done by our excellent collaborators at Cambridge (Payam Gammage, Lindsey Van Haute and Michal Minczuk). We first expand a mathematical model for how diverse mtDNA populations within cells change over time – building new power stations and decommissioning old ones, under the “governance” of the cell. We then produce an “energy budget” for the cellular “society” – describing the costs of building, decommissioning, and maintaining different power stations, and the corresponding profits of energy generation.

We find some surprising results. First, it can get harder to maintain a good energy budget in a tissue (a collection of individual cellular “societies”) over time, even if demands stay the same and average mtDNA quality doesn’t change. This is because the cell-to-cell variability in mtDNA quality does increase, carrying with it an added energetic challenge. This increased challenge could be a contributing factor to the collection of problems involved in ageing.

An overview of our approach. A mathematical model for the processes and "budget" involved in controlling mtDNA populations makes a general set of biological predictions and explains gene-therapy observations

Next, we found that cells with only low-quality mtDNA can perform worse than cells with a mix of low- and high-quality mtDNA. This is because low-quality mtDNA may consume less cellular resource, although global efficiency is decreased. Linked to this, removal of low-quality mtDNA (decommissioning bad power stations) alone is not always the best strategy to improve performance. Instead, jointly elevating low- and high-quality mtDNA levels, avoiding this detrimental mixed regime, is the best strategy for some situations. These insights may help explain some of the negative effects recently observed in cells with mixed mtDNA populations.


Our theory suggests that mixed mtDNA populations may do worse than pure ones, even if the pure population is a low-functionality mutant. Image from Hanne's post here 


We identified how best to control cellular mtDNA populations across the full range of possible populations, and used this insight to link with exciting gene therapies where low-quality mtDNA is preferentially removed through an experimental intervention (using so-called “endonucleases” to cut particular mtDNA molecules). We found that strong, single treatments will be outperformed by weaker, longer-term treatments, and identified how the mtDNA variability we know is present can practically effect the outcome of these therapies. We hope that the principles found in this work both add to our basic understanding of ageing and mixed (“heteroplasmic”) mitochondrial populations, and may inform more efficient therapeutic approaches in the future. Iain, Hanne, Nick

Thursday, 11 July 2019

ARTICLE: Coupling mitochondrial physics and genetics

Mitochondrial Network State Scales mtDNA Genetic Dynamics
Juvid Aryaman, Charlotte Bowles, Nick S. Jones and Iain G. Johnston
Genetics Early online July 10, 2019; https://doi.org/10.1534/genetics.119.302423

Mitochondrial DNA (mtDNA) populations within our cells encode vital energetic machinery. MtDNA is housed within mitochondria, cellular compartments lined by two membranes, that lead a very dynamic life. Individual mitochondria can fuse when they meet, and fused mitochondria can fragment to become individual smaller mitochondria, all the while moving throughout the cell. The reasons for this dynamic activity remain unclear (we’ve compared hypotheses about them before here and here, with blog articles here). But what influence do these physical mitochondrial dynamics have on the genetic composition of mtDNA populations?

MtDNA populations can, naturally or as a result of gene therapies, consist of a mixture of different mtDNA types. Typically, different cells will have different proportions of, say, type A and type B. For example, one cell may be 20% type A, another cell may be 40% type A, and a third may be 70% type A. This variability matters because when a certain threshold (often around 60%) is crossed for some mtDNA types, we get devastating diseases.

We previously showed mathematically (blog) and experimentally (blog) that this cell-to-cell variability in mtDNA proportions (often called “heteroplasmy variance” and sometimes referred to via the “mtDNA bottleneck”) is expected to increase linearly over time. However, this analysis pictured mtDNAs as individual molecules, outside of their mitochondrial compartments. When mitochondria fuse to form larger compartments, their mtDNA is more protected: smaller mitochondria (and their internal mtDNA) are subject to greater degradation. More degradation means more replication, and more opportunities for the fraction of a particular type of mtDNA to change per unit time. In a new paper here in Genetics, we show (using a mathematical tour de force by Juvid) that this protection can dramatically influence cell-to-cell mtDNA variability. Specifically, the rate of heteroplasmy variance increase is scaled by the proportion of mitochondria that exist in a fragmented state. (It turns out that it's the proportion of itochondria that are fragmented that's important -- not whether the rate of fission-fusion is fast or slow).


This has knock-on effects for how the cell can best get rid of low-quality mutant mtDNA. In particular, if mitochondria are allowed to fuse based on their quality (“selective fusion”), we show that intermediate rates of fusion are best for removing mutants. Too much fusion, and all mtDNA is protected; too little, and good mtDNA cannot be sorted from bad mtDNA using the mitochondrial network. This mechanism could help explain why we see different levels of mitochondrial fusion in different conditions. More broadly, this link between mitochondrial physics and genetics (which we’ve also speculated about here (blog) and here) suggests one way that selective pressures and tradeoffs could influence mitochondrial dynamics, giving rise to the wide variety of behaviours that remain unexplained. Juvid, Nick, and Iain

ARTICLE: Getting to the root of the problem

Model selection and parameter estimation for root architecture models using likelihood-free inference
Clare Ziegler, Rosemary J. Dyson, Iain G. Johnston
J Roy Soc Interface (online, doi.org/10.1098/rsif.2019.0293 , 2019)

Roots bridge plants and soil, making vital contributions to crops, the environment, and fundamental biology. Because of this importance, understanding how roots grow under different conditions is a key scientific target. Experimental approaches to study roots can be challenging: being underground, it’s hard to observe root systems without perturbing them. Computer models can help here: we can simulate root growth and the “architecture” of root systems under lots of different conditions, without having to dig up and destroy real plants.


 
Observing roots growing underground is hard, but not impossible: here's a shot from our "minirhizotron" experiments using underground cameras to watch roots grow in an experimental woodland facility (see article here, and 3D version here!)

As computers have become more powerful, more and more sophisticated models for root growth and architecture have emerged. These simulation approaches typically take as input a set of parameters, and produce as output a model root system. These parameters are numbers describing, for example, the rates of root elongation, distances between lateral root branches, and so on – there may be dozens, or hundreds, of parameters in a sophisticated root model.

The output of a model depends strongly on these parameter values. So how can we choose the “right” ones? We may know some from experiments – for example, the widths of roots can be readily measured. But others may be less easy to observe. It is quite common to make educated guesses at these parameters, and see if the resulting root system “looks right”. This approach has a few issues – it can be subjective, and doesn’t give us information on how flexible our guesses are. For example, is a growth rate of 0.1cm per day just as likely as 0.5cm per day, or 0.02cm per day? And how can we tell if one version of a model does “better” than another, and is more supported by real observations?

In a new paper here in Journal of the Royal Society Interface, we propose a platform to provide answers to these questions, using so-called “approximate Bayesian computation” or ABC. This is a way of learning which parameter values and models are most compatible with observed data, by running many simulations with many different choices, and comparing the output of each choice to our observations using specific criteria. This replaces the subjective “looks right” and explores a wide set of parameterisations, allowing us to learn what ranges of values are most likely given our data. We can also use ABC to compare different mechanisms for root growth, finding which is most supported by observation. This helps us gain scientific insight and ensures that the outputs of our models can be more reliably intepreted.


Overview of our approach. Using ABC allows us to identify governing parameters and mechanisms for root growth that are most supported by real observations.

We tested our ABC approach with synthetic observations from models of thale cress and narrowleaf lupin, confirming that we can recover the parameter values we put in. We then used real thale cress plants (wild and mutant) to show that our platform distinguishes genetically different plants and identifies most-likely parameters and model structures for real root growth. We used the platform to select models for growth and branching, showing how it can be used to compare existing models from the literature. We hope that this approach can be used to further help improve the interpretability and rigour of plant modelling and simulation! Iain and Clare

Saturday, 22 September 2018

ARTICLE: Time marches on -- mitochondria, ageing, and disease

Burgstaller, J.P., Kolbe, T., Havlicek, V., Hembach, S., Poulton, J., Piálek, J., Steinborn, R., Rülicke, T., Brem, G., Jones, N.S. and Johnston, I.G. Large-scale genetic analysis reveals mammalian mtDNA heteroplasmy dynamics and variance increase through lifetimes and generations. Nature communications2488 (2018)

DNA in mitochondria, the powerhouses of the cell, is passed down from mother to child. But there are many mitochondria in each cell, and these mitochondria may have different genetic features. If a mother carries a mixture of mitochondrial DNA (mtDNA) types, this can make it hard to say which features their children will inherit. For mothers carrying a disease-causing mtDNA mutation, this makes family planning and clinical therapies challenging.


In particular, the role of a mother's age has long been a mystery. Is the probability of a child inheriting a particular mtDNA feature higher when mothers are younger or older? An answer to this question could help plan clinical strategies to improve fertility and prevent the inheritance of deadly mitochondrial disease.


To address this, we worked with our excellent collaborators with a combination of maths, statistics, and experiment. Our collaborators used cutting-edge technology to reveal the mixtures of mtDNA in the egg cells of mother mice at a wide range of ages, and in the litters of offspring the mothers produced. This experimental work was the largest-scale study of mammalian mtDNA that we're aware of, involving thousands of observations throughout lifetimes and between generations. In concert, we developed a mathematical model describing the changes to, and inheritance of, mtDNA from mother to offspring. We combined the model and data to learn how different biological processes affect mtDNA through and between generations.



Cells contain populations of mitochondria, and these populations change over time. In European mice, we observed how variability in these populations evolves as mammals age and reproduce. We found that older mother have more varied mitochondria and pass this variance on to their offspring -- of central importance in the inheritance of genetic disease. 

We found that the variability of mtDNA dramatically increased as mothers aged. This means that the probability of inheriting more extreme -- both lower and higher -- levels of a genetic feature increases for older mothers. We also found that different mtDNA mixtures were inherited in different ways - with some mtDNA types favoured for inheritance and some disfavoured. We used our findings to create a way to predict how the risk that offspring would inherit disease-causing mtDNA features changes over time. Moving forward, we're aiming to harness these powerful ways of using large datasets to describe and predict the dynamics of mtDNA inheritance in humans, and to learn what it is about these mtDNA types that predicts their evolution across generations. You can read the article for free in Nature Communications here.

ARTICLE: How do plants roll dice?

Johnston, I.G. and Bassel, G.W. Identification of a bet-hedging network motif generating noise in hormone concentrations and germination propensity in Arabidopsis. Journal of the Royal Society Interface15 141 (2018)

Seeds feed the world, and uniform, reliable harvests of seeds and grains is essential for food security. However, there's a fundamental tension between the evolutionary priorities of plants and the agricultural priorities of humans. Evolutionarily, it is good for plants to "hedge their bets" by having seeds germinate at different times. A plant whose seeds all germinate in March will be susceptible to a frost in April, potentially leading to the loss of a generation of offspring. By contrast, a plant whose seeds germinate throughout March and April will have a subset of its offspring survive that frost, and its genes will be passed on to the next generation.


This bet-hedging poses a challenge for agriculture. In agricultural settings, we have more control over plant environments, and so plants have less need to withstand unpredictable environmental fluctuations. At the same time, non-uniform germination decreases crop yields, makes harvesting harder, and makes crops more susceptible to pest invasion. If we can learn how plants generate this evolved germination variability, we can design engineering and/or breeding strategies to reduce this and improve crop yields.



Plants have evolved to "hedge their bets" by having seeds germinate at different times -- this makes generations of plants more robust to environmental fluctuations. Our work reveals a mechanism that "rolls dice" within plant cells, acting like a random number generator to produce variability in germination propensity. 

In a previous paper (blog post here), we looked at how germination is controlled by an interaction between two hormones known as ABA and GA. During that project, we noticed a surprising feature of the cellular pathways affecting ABA. Oddly, it seemed that ABA both activated a pathway that increased its own production, and at the same time (and in the same place) activated a pathways that increased its own degradation. These two pathways seemed to be competitive -- one increases levels of ABA, the other decreases them. Why would cells spend energy in this "futile" way?


We hypothesised that these competitive pathways might have the effect of generating variability in ABA levels. The pathways are fundamentally "noisy", involving random interactions in the chaotic environment of the cell. Consider increasing the activity of both pathways simultaneously. One pathway would act to increase levels of ABA, the other would act to decrease it. The increased "push and pull" of these noisy pathways would increase the spread of levels of ABA in different cells, even if average levels stayed the same.


Because it's hard to measure the levels of hormones in individual cells over time, we initially took a theoretical approach. We showed, with maths, that the competing pathways did indeed have this variability-inducing effect. By varying the activity through these pathways, the cell can increase variability in ABA levels, and hence increase variability in germination propensity. We showed that the theory we developed was compatible with some experiments where the ABA circuitry was artificially manipulated. The theory went on to reveal various aspects of cellular machinery that we could conceivably target through synthetic approaches, in order to reduce germination variability. Put together, our quantitative theory, supported by experiment, explained the mysterious competitive pathways and revealed several new interventions with the potential to improve food security. You can read about it for free in the Journal of the Royal Society Interface here. Iain  


ARTICLE: Which genes are essential for bacterial survival?

Goodall, E.C., Robinson, A., Johnston, I.G., Jabbari, S., Turner, K.A., Cunningham, A.F., Lund, P.A., Cole, J.A. and Henderson, I.R., 2018. The essential genome of Escherichia coli K-12. mBioe02096 (2018)

Bacteria cause diseases, and are developing resistance to the drugs we use to kill them. Anti-microbial resistance (AMR) is one of the most pressing global health challenges facing society. In the immense scientific endeavour of creating new, effective treatments for bacterial infections, fundamental biological knowledge about how bacteria live and proliferate is of vital importance.


One way we can obtain this knowledge is by discovering what cellular machinery that bacteria need to survive and proliferate. A common (and famous) bacterium called Escherichia coli (E. coli) has over 4000 protein-coding genes, but we're not really sure which of these genes is essential for the bacterium, and how many provide some non-essential "added value". If we can learn which genes are essential for bacteria, we have a more specific set of targets to shoot for in designing new drugs and therapies.


So -- how can we find out which genes are essential for E. coli? One neat way involves a new experimental approach called transposon-directed insertion site sequencing (TraDIS). Transposons are elements of DNA that can be inserted into a bacterial genome -- when they are inserted into part of the genome that codes for a gene, they prevent that gene being properly expressed, effectively removing it from the bacterium. TraDIS, in essence, takes a large population of bacteria and inserts one transposon into a random position in each bacterium. The population is then left to evolve for some time. After that time, we look at the genomes of bacteria within the surviving population, and see exactly where transposon insertions have been retained in some living bacteria.



A stylised representation of the E. coli genome and the positions within it where we found transposons to have been retained (corresponding to non-essential genes). 

The idea is that any bacteria in the population that have a transposon inserted into an essential gene will die. As such a gene is essential, it's required for survival, and a transposon preventing its expression will kill the bacterium. Therefore, if some bacteria in a population retain an insertion in gene X and survive, it follows that gene X is not essential. Conversely, if we see a large region of the genome within which no insertions are retained in the final population, it is likely that that region corresponds to an essential gene. 


There's some mathematical subtlety in the "it is likely". Depending on how many transposon insertions originally occur, and the length of the genome, some regions without insertions may occur just by chance. We did a bit of maths to work out how unlikely it is to see an insertion-free region of a given length arise by chance; and, by extension, how likely it is that a gene identified by this analysis is indeed essential for the bacterium. However, the maths was only one part of this project -- it was first and foremost an experimental tour de force by our excellent collaborators. We jointly provided a new atlas of essential genes in E. coli, provide a new way of reasoning about the powerful TraDIS technique, and provide several new insights into bacterial physiology and biochemistry. The work is freely available in the journal mBio here. Iain 

ARTICLE: How plants decide when to germinate

Topham, A.T., Taylor, R.E., Yan, D., Nambara, E., Johnston, I.G. and Bassel, G.W. Temperature variability is integrated by a spatially embedded decision-making center to break dormancy in Arabidopsis seeds. PNAS 114 6629 (2017)

A plant's choice to germinate is one of the most important decisions in the world. If it is made too soon, the plant may be damaged by harsh winter conditions; if too late, the plant may be outcompeted, and crop yields may be lower. If crops in a field make the decision at different times, there is more room for weeds to grow and pests to take over. 


In a recent study, we combined mathematical modelling with several neat experiments to identify sets of cells that make this germination choice in a much-studied plant called thale cress (Arabidopsis thaliana), and have learned how it makes decisions based on the plant's environment.



Two views of the plant embryo from laser microscopy, highlighting cells where different components of the germination control machinery are expressed. The background shows the "attractor basins" in a mathematical description of the germination decision: horizontal and vertical axes give the levels of two hormones ABA and GA, the blue region corresponds to dormant seeds and the red region to germination. 

This germination circuitry functions through a circuit of chemical stimuli and responses. Using laser microscopy, we found that different parts of this circuit exist in different parts of the plant embryo -- and that the separation of these parts is central to how the brain functions. We used mathematical modelling to show that communication between separated elements of the germination circuitry controls the plant's sensitivity to its environment. Following this theory, we used a mutant plant where cells were more chemically linked -- essentially enhancing communication between circuit elements -- to show that germination depends on these intra-cellular signals.


The separation of circuit elements allows a wider palette of responses to stimuli. It's like the difference between reading one critic's review of a film four times over, or amalgamating four different critics' views before deciding to go to the cinema. Our mathematical theory predicted that more plants would germinate when exposed to varying environments -- like three short pulses of cold -- than constant environments -- like one long cold period. We tested this theory in the lab and found exactly this behaviour.


Next, the hope is to learn about the germination brain in other plants and crops, and to show how our new knowledge of the germination machinery can be used to enhance and synchronise germination in crops. You can read the paper for free in the journal PNAS here. Iain

Sunday, 21 May 2017

ARTICLE: A healthy dose of mathematics

Toward Precision Healthcare: Context and Mathematical Challenges
C Colijn, N Jones, IG Johnston, S Yaliraki, M Barahona
Frontiers in Physiology 8 136 (2017)
  • The continuing explosion of available biomedical data will help us tailor and optimise therapies for individual patients; we are designing new maths and statistics to help this process and to include social and other data into an overarching "precision healthcare" approach.
Our research combines tools from maths and statistics with biological data to learn more about the biological world. An exciting, growing, and much-discussed branch of science -- precision medicine -- is a specific instance of this idea. The vision of precision medicine is to use the expanding volume of data that's emerging from medicine and biology to tailor and optimise medical therapies for individual patients, making the therapies as effective as possible. This idea isn't new -- we are well aware, for example, that an individual's blood type dictates which blood transfusions they can successfully receive. But precision medicine is a much bigger picture, potentially taking into account large amounts of genetic, environmental, dietary, and other features to identify the optimal treatment for a disease -- for example, tailoring chemotherapy treatments to match the genetic specifics of a particular cancer case.

Dealing with these large and diverse datasets will need new mathematical and statistical approaches, built with an ongoing link to clinical practice. At the same time, we're interested in expanding the idea of precision medicine to include the "big data" that's increasingly available about individuals' social and logistic contexts. Social networks can dictate how diseases spread -- and how knowledge and views about therapies, vaccines, and other medically pertinent ideas are transmitted and shaped from person to person. A person's home region determines the genetic structure of local people who may act as donors. We're looking at the idea of "precision healthcare" -- using new maths and statistics to optimise healthcare strategy, not just individual therapies, in the light of large-scale datasets.





One aspect of precision healthcare we'll be exploring is exploring how progressive diseases -- those that involve the accumulation of symptoms over time -- develop in patients, using transition networks like those above to model "disease spaces" and find pathways in those spaces.

We're excited to be part of a new initiative -- the Centre for the Mathematics of Precision Healthcare -- involving six parallel and related research projects that align with this goal. Some of our previous work -- for example, estimating social structures of big UK cities to explore the challenges that genetic diversity poses to gene therapies for mtDNA disease -- already has a precision healthcare feel. In a new review paper (available for free) we discuss this and other examples of past and future work that we hope will contribute to the precision healthcare goal, along with some key ideas and context for the initiative. Iain

Monday, 31 October 2016

ARTICLE: The maths of mitochondrial DNA

Evolution of Cell-to-Cell Variability in Stochastic, Controlled, Heteroplasmic mtDNA Populations
IG Johnston, NS Jones
The American Journal of Human Genetics 99 (5), 1150-1162 (2016)
  • Vital populations of mtDNA are constantly evolving in our cells in response to random influences and control from the nucleus: we build a general mathematical theory describing this poorly-understood process and show that it predicts a wide range of existing experimental outcomes and gives us lots of new insights into biology and disease
Mitochondrial DNA (mtDNA) contains instructions for building important cellular machines. We have populations of mtDNA inside each of our cells -- almost like a population of animals in an ecosystem. Indeed, mitochondria were originally independent organisms, that billions of years ago were engulfed by our ancestor's cells and survived -- so the picture of mtDNA as a population of critters living inside our cells has evolutionary precedent! MtDNA molecules replicate and degrade in our cells in response to signals passed back and forth between mitochondria and the nucleus (the cell's "control tower"). Describing the behaviour of these population given the random, noisy environment of the cell, the fact that cells divide, and the complicated nuclear signals governing mtDNA populations, is challenging. At the same time, experiments looking in detail at mtDNA inside cells are difficult -- so predictive theoretical descriptions of these populations are highly valuable.

Why should we care about these cellular populations? MtDNA can become mutated, wrecking the instructions for building machines. If a high enough proportion of mtDNAs in a cell are mutated, our cells struggle and we get diseases. It only takes a few cells exceeding this "threshold" to cause problems -- so understanding the cell-to-cell distribution of mtDNA is medically important (as well as biologically fascinating). Simple mathematical approaches typically describe only average behaviours -- we need to describe the variability in mtDNA populations too. And for that, we need to account for the random effects that influence them.
 

In our cells, signals from the "control tower" nucleus lead to the replication (orange) and degradation (purple) of mtDNA. These processes affect mtDNA populations that may contain normal (blue) and mutant (red) molecules. Our mathematical approach -- extending work addressing a similar but simpler system -- describes how the total number of machines, and the proportion of mutants, is likely to behave and change with time and as cells divide.

In the past, we have used a branch of maths called stochastic processes to answer questions about the random behaviour of mtDNA populations. But these previous approaches cannot account for the "control tower" -- the nucleus' control of mtDNA. To address this, we've developed a mathematical tradeoff -- we make a particular assumption (which we show not to be unreasonable) and in exchange are able to derive a wealth of results about mtDNA behaviour under all sorts of different nuclear control signals. Technically, we use a rather magical-sounding tool called "Van Kampen's system size expansion" to approximate mtDNA behaviour, then explore how the resulting equations behave as time progresses and cells divide.

Our approach shows that the cell-to-cell variability in heteroplasmy (the potentially damaging proportion of mutants in a cell) generally increases with time, and surprisingly does so in the same way regardless of how the control tower signals the population. We're able to update a decades-old and commonly-used expression (often called the Wright formula) for describing heteroplasmy variance, so that the formula, instead of being rather abstract and hard to interpret, is directly linked to real biological quantities. We also show that control tower attempts to decrease mutant mtDNA can induce more variability in the remaining "normal" mtDNA population. We link these and other results to biological applications, and show that our approach unifies and generalises many previous models and treatments of mtDNA -- providing a consistent and powerful theoretical platform with which to understand cellular mtDNA populations. The article is in the American Journal of Human Genetics here and a preprint version can be viewed here. Iain

Friday, 28 October 2016

ARTICLE: Random number seed

Variability in seeds: biological, ecological, and agricultural implications 
J Mitchell, IG Johnston, GW Bassel
Journal of Experimental Botany, erw397 (2016) 
  • Natural variability across scales, from the molecular to the environmental, means that individual seeds behave differently; we explore the challenges this variability poses for agriculture and food security, and how modern science can help address these challenges.
Seeds feed the world. Whether eaten themselves, or allowed to develop into crop plants which are then consumed by humans or livestock, seeds are the fundamental starting point for agriculture. But each seed has a different story. Throughout millions of years of evolution, plants have evolved to -- forgive the pun -- "hedge" their bets from one generation to the next. A parent plant cannot completely predict the environmental conditions that its offspring will face, so it induces variability in the seeds it produces. If some seeds are better at surviving in environment A and some are better in environment B, the plant has a way of ensuring its genes will survive regardless of whether the environment is A-like or B-like in future.

This bet-hedging is a sensible evolutionary strategy when environments are unpredictable. But modern agriculture makes environments much more predictable than the wild situations plants have been exposed to throughout evolutionary history. Now bet-hedging becomes a bad thing -- if we know the environment will always be C, energy spent ensuring that seeds survive in environments A and B is wasted, reducing potential yields.

Understanding and controlling the variability within populations of seeds thus has huge implications for agriculture. Variability inherent within populations of seeds, in addition to differences in the environments that seeds experience, means that, for example, seed lots germinate asynchronously (some quickly, some slowly or not at all). This leads to non-uniform and sub-optimal crop production, allows pests to enter fields, and challenges our ability to plan agricultural strategies. If we could control seed variability, these problems would be diminished, with a host of positive consequences for food security.

A given set of seeds will vary in their behaviour due to influences on many scales, from random molecular processes within cells to large-scale environmental stimuli. As a result, important features like germination propensity vary across seed lots (perhaps taking a broad distribution like that illustrated here), posing a challenge to agriculture and food security, which scientific understanding can mitigate.

In a new review, we survey our current understanding of the sources of variability in seeds, and its biological and agricultural implications. Processes across many scales induce variability in seed behaviour, from random cell biological interactions (like we've written about before!), through seed position in a parent plant, to large-scale environmental differences. We particularly focus on germination, an aspect of seed behaviour of crucial biological and agronomic importance, which takes place when a "developmental switch" in a seed is flipped. We discuss the genetic and molecular players that modern science has discovered to influence this decision to germinate in seeds, and describe the challenges in furthering our understanding of this vital question -- and how cool new tech, and maths, can help us make new progress! The review is in the Journal of Experimental Botany here. Iain