Showing posts with label model selection. Show all posts
Showing posts with label model selection. Show all posts

Wednesday, 23 February 2022

ARTICLE: Social networks of plant mitochondria

Network analysis of Arabidopsis mitochondrial dynamics reveals a resolved tradeoff between physical distribution and social connectivity
Joanna M Chustecki, Daniel J Gibbs, George W Bassel, Iain G Johnston
Cell Systems 12 419 (2021)

We recently spent some time looking at a long-standing question in plant cell biology -- why do mitochondria move the way they do? Plant mitos look for all the world like cars in a city, moving along highways and speedily getting from place to place. We combined laser microscopy, video analysis, physical modelling, and network science to explore what benefits this motion might bring to the cell. It turns out, it allows mitochondria to have social lives! Through the "social network" of encounters between moving mitochondria, beneficial exchange of contents can occur, while their motion allows the cell to keep its population well spread. Here's a blog article from Jo explaining things more!

Read more...

... and also check out Jo's beautiful site!

Mitochondria are in yellow in the microscopy image; their "social network", describing their encounters, is overlaid in white. Cover of the month's Cell Systems issue.

ARTICLE: Corals to crops -- how life protects the plans for its cellular power stations

Avoiding organelle mutational meltdown across eukaryotes with or without a germline bottleneck
David M Edwards, Ellen C Røyrvik, Joanna M Chustecki, Konstantinos Giannakis, Robert C Glastad, Arunas L Radzvilavicius, Iain G Johnston
PLoS Biology 19 e3001153 (2021)

(this text is from a press release about the article)

An international team of researchers led by the University of Bergen has uncovered how organisms from crops to corals may avoid deadly DNA damage during evolution.

Our cells, and those of animals, plants and fungi, contain compartments that produce chemical fuel. These compartments contain their own DNA, which stores instructions for important cellular machinery. But this so-called oDNA (organelle DNA) can become mutated, corrupting the instructions and preventing cells making enough energy.

In humans and some other animals, a process called the “bottleneck” allows some offspring to inherit less mutated oDNA. This process needs mothers’ egg cells to develop early, like in humans, where a human girl is born with all her egg cells already formed. But other organisms, from plants to fungi, don’t develop these cells early – their flexible body plans mean that eggs are not “set aside” early in development.<\p>

“We wanted to know how these organisms might avoid inheriting mutations without a human-like bottleneck,” said Ellen Røyrvik, a geneticist on the research team, based at UiB.

The scientists used mathematical modelling to show that a process called gene conversion – the controlled overwriting of DNA – could in theory allow some offspring to inherit less mutant oDNA without requiring a bottleneck. Using genome data, they found machinery controlling this process in plants and fungi, but also in soft corals, sponges, and algae – all organisms without fixed body plans. They also found that this machinery was most active in the parts of plants that will end up producing the seeds of the next generation, suggesting that it is indeed used to allow some offspring to inherit fewer mutations.

Organisms without fixed body plans (including octocorals, sea pens, sponges, plants, and fungi) and with fixed body plans (including humans and many animals) may use different strategies to avoid the buildup of damage in their cellular "power stations." CREDIT: Gemma Lofthouse

“Taken together, it looks like organisms without a fixed body plan – plants, fungi, corals, sponges, algae – may have adopted gene conversion to deal with oDNA mutations,” said Iain Johnston, an associate professor in the Mathematics Institute at UiB, who led the research. “Humans and other animals can develop egg cells early and use a bottleneck; other organisms can use gene conversion instead.”

Going forward, the team plans to explore how this overwriting of oDNA causes other issues in the organisms that use it – including crop plants, where it can cause sterility. They are also exploring the broader question of why these compartments contain oDNA at all, given the risk of mutational damage.

Friday, 23 April 2021

ARTICLE: What makes mitochondria selfish, and when do selfish ones win?

MtDNA sequence features associated with ‘selfish genomes’ predict tissue-specific segregation and reversion, Nucleic Acids Research 48 8290 (202)

Mitochondria, the power stations of the cell, are in some senses like people in a company. The company needs several people to contribute if it is to survive. Some people may work hard and contribute lots to the company. Others may selfishly slack off and rely on others doing the work.

The cell needs mitochondria to produce ATP, the chemical that powers many important processes. But there is evidence that some mitochondria are more selfish, and some less so, than others. Unselfish mitochondria produce machinery which helps produce ATP. Selfish mitochondria prefer to replicate, copying their DNA and contributing less to the cell. Interestingly, at the molecular level, there is something like a "switch": a mitochondrion either takes steps to produce useful machinery, or takes steps that will help it replicate.

We were interested in why different mitochondria choose different positions of this switch, and what is a good "strategy" for mitochondria under different conditions. We built a simple model of this behaviour to understand it. Unsurprisingly, we found that selfish mitochondria -- favouring replication, and contributing less to the cell -- profilerate over unselfish ones in cells where there's little pressure to co-operate. As they replicate more, selfish mitochondria eventually come to dominate such cells. Where there is cellular pressure, however, unselfish mitochondria may win out. This is because cells full of selfish mitochondria won't perform adequately, and the whole cell and all its mitochondria will die -- leaving those cells with more unselfish mitochondria remaining.

What do we mean by "cellular pressure"? Well, if some type of cells never die, clearly the latter event can't happen, and we'd expect selfish mitochondria to win. If cells die regularly, perhaps there's more capacity to select those filled with unselfish mitochondria. We looked at different tissues where cells die with different rates, in mice where cells had two different types of mitochondrial DNA (mtDNA). We found a consistent pattern where one type of mtDNA proliferated in slow-dying cells and the other proliferated in fast-dying cells. Why different mtDNA types "win" in different tissues is a big question (which we've looked at before!), and it looks like this might help explain some of these differences.


(A) Sequence features may make different mtDNA types more "selfish" (favouring replication) or "unselfish" (favouring the production of useful machinery). (B) Our theory shows how, depending on cellular pressures, one or the other strategy can be favoured, leading to selection for one or the other mtDNA type.

We also asked what it is about a particular mtDNA sequence that might make it more or less selfish. Based on how mtDNA produces useful machinery, and how it replicates, we hypothesised that some features in the so-called "control region" of mtDNA may influence selfishness. Using sequence information, we found that these features tied quite neatly in with the observations in these mouse models, and also in (more limited) observations from human cells. While certainly not resolved, this picture suggests a link between sequence features of mtDNA, cellular selfishness, and proliferation differences across different tissues. You can read more (for free) in Nucleic Acids Research here.

Monday, 8 June 2020

ARTICLE: Transport planning in biology

Efficient vasculature investment in tissues can be determined without global information
S Duran-Nebreda, IG Johnston, GW Bassel
Journal of the Royal Society Interface 17 20200137 (2020)


We need roads. Roads link up different parts of our society, allowing us to send messages and supplies from one region to another. But they come at a cost. If we lay down a road across the country, we can't use that land to farm or build houses, and maintaining roads costs a lot of tax money.

Multicellular organisms have the same issue. They also need to send supplies (e.g. nutrients) and messages (e.g. chemical signals) from one place to another. So they build roads. Our blood vessels are one example, transporting oxygen and hormonal messages throughout our bodies. So-called vasculature -- our blood vessels are one example, as are xylem and phloem in plants -- is used to allow transport around an organism. But again, if some parts of the organism are being used for transport, they can't be used for doing other useful things.

Given this cost of producing "roads", organisms would presumably like to be efficient as possible when laying out their transport systems. This may involve, for example, making journey lengths as short as possible while using as little land as possible for roads. But while city planners and engineers can look at maps and run simulations to work out how best to place roads, organisms lack a top-down "planner" with a large-scale map. How then do organisms efficiently resolve this tradeoff? Specifically, how is it decided where best to place vasculature to minimise the effective distance between cells?

We took a look at this using a theoretical model where an organism's tissue is modelled as a collection of cells in a 2D layer, a 3D block, or an intermediate case involving a set of layers, or a more realistic structure taken from experimental characterisation of plant tissues. We considered different ways that an organism might produce vasculature by fusing together cells in this model tissue to make "roads". This method for making vasculature models the case in immobilised cells, like we find in plants. We considered different ways that cells might be chosen to fuse, based on the physical structure of the tissue, and allowing some randomness in this decision.


 How has this plant made efficient "roads" (vasculature, like the veins seen here) without having a map of the whole leaf? We found that it can do a pretty good job without a global map, just using local sensing.

We found that using a "top-down" planner (with a map of all cells – which organisms don't have!) to choose which cells to fuse is usually the best way of producing an efficient transport network. But, we found that "bottom-up" approaches, where cells fuse based on purely local information (as opposed to a global map of the whole tissue) can actually do almost as well as the top-down planner. Strikingly, we found that these bottom-up approaches can provide "scale-free" improvements in transport. This means that the amount by which having more roads decreases journey lengths doesn't depend on the overall size of the system. The transport improvements from vasculature were more pronounced in 3D than in 2D, and the best approach for vasculature production varied in the different plant tissues we looked at. This suggests that there may be some evolutionary back-and-forth between the rules that plants use to create vasculature and the form of their tissues, which we plan to explore further in future!

Wednesday, 8 January 2020

ARTICLE: Learning pathways of disease progression

HyperTraPS: Inferring probabilistic patterns of trait acquisition in evolutionary and disease progression pathways
Sam F Greenbury, Mauricio Barahona, Iain G Johnston
Cell Systems (2019)

Many diseases that take a substantial human toll can be viewed as “progressive”. That is, a patient starts out healthy, then disease-related problems and/or symptoms develop over time. For example, a given case of cancer may begin with a patient acquiring a particular mutation, then other mutations building up in their genome over time.

How the same disease progresses in different patients often varies widely. Understanding this variability is important for precision medicine, where detailed knowledge of individual patients is used to design the best targeted treatments. However, learning the varied pathways of diseases and using them to predict future outcomes is challenging. Human researchers usually cannot hope to remember or analyse enough examples of patient data to provide the most reliable picture.

We previously developed an algorithm called HyperTraPS (hypercubic transition path sampling) to explore how biological systems evolve over time. We reasoned that HyperTraPS could also be used to learn the pathways of disease progression. In a new study in Cell Systems (free preprint available here) we used HyperTraPS to analyse biomedical data from many patients – hundreds, or thousands of individuals – to build a ‘road map’ of the different pathways that a disease takes over time.

Picture a river that branches out into a wide delta. Patients start out healthy – upstream in the river – and different patients go down different branches as the disease progresses and they acquire more symptoms. HyperTraPS learns the structure of the river delta from data, and predicts which river branches are more or less likely – and, importantly, where you'll end up if you're currently at a particular point.

By learning these branching patterns of disease progression, HyperTraPS has helped provide a refined risk assessment for malaria, based on data from thousands of Gambian children – as we’ve written about before. The approach also revealed diverse pathways of ovarian cancer progression, where the first mutation to occur appears to play a large role in determining subsequent mutations.

The "waterfall" in the foreground shows paths from one stage of a disease to the next, learnt by HyperTraPS using data from a high number of patients. Each dot of the illustration represents different stages of disease, for example a specific set of symptoms or a given set of mutations. The thickness of the lines indicate the probability of moving from one specific stage of disease to the next.

HyperTraPS is very generalisable and can be used to learn pathways by which mutations, symptoms, or other features develop over time from an initial state. We further used this generalisability to understand a biomedically important example of evolution – specifically, how tuberculosis evolves to become resistant to antibiotics.

Tuberculosis acquires resistance through mutations, and HyperTraPS has revealed the patterns of these mutations in TB bacteria reported from a group of 1000 Russian patients. These patterns help predict which mutation a bacterium will acquire next, and hence which drugs may be more effective for a given case. We’re following up with other applications of HyperTraPS, to learn about other progressive diseases, ageing, and evolution, and even to analyse how students complete tasks in online courses.

Thursday, 11 July 2019

ARTICLE: Getting to the root of the problem

Model selection and parameter estimation for root architecture models using likelihood-free inference
Clare Ziegler, Rosemary J. Dyson, Iain G. Johnston
J Roy Soc Interface (online, doi.org/10.1098/rsif.2019.0293 , 2019)

Roots bridge plants and soil, making vital contributions to crops, the environment, and fundamental biology. Because of this importance, understanding how roots grow under different conditions is a key scientific target. Experimental approaches to study roots can be challenging: being underground, it’s hard to observe root systems without perturbing them. Computer models can help here: we can simulate root growth and the “architecture” of root systems under lots of different conditions, without having to dig up and destroy real plants.


 
Observing roots growing underground is hard, but not impossible: here's a shot from our "minirhizotron" experiments using underground cameras to watch roots grow in an experimental woodland facility (see article here, and 3D version here!)

As computers have become more powerful, more and more sophisticated models for root growth and architecture have emerged. These simulation approaches typically take as input a set of parameters, and produce as output a model root system. These parameters are numbers describing, for example, the rates of root elongation, distances between lateral root branches, and so on – there may be dozens, or hundreds, of parameters in a sophisticated root model.

The output of a model depends strongly on these parameter values. So how can we choose the “right” ones? We may know some from experiments – for example, the widths of roots can be readily measured. But others may be less easy to observe. It is quite common to make educated guesses at these parameters, and see if the resulting root system “looks right”. This approach has a few issues – it can be subjective, and doesn’t give us information on how flexible our guesses are. For example, is a growth rate of 0.1cm per day just as likely as 0.5cm per day, or 0.02cm per day? And how can we tell if one version of a model does “better” than another, and is more supported by real observations?

In a new paper here in Journal of the Royal Society Interface, we propose a platform to provide answers to these questions, using so-called “approximate Bayesian computation” or ABC. This is a way of learning which parameter values and models are most compatible with observed data, by running many simulations with many different choices, and comparing the output of each choice to our observations using specific criteria. This replaces the subjective “looks right” and explores a wide set of parameterisations, allowing us to learn what ranges of values are most likely given our data. We can also use ABC to compare different mechanisms for root growth, finding which is most supported by observation. This helps us gain scientific insight and ensures that the outputs of our models can be more reliably intepreted.


Overview of our approach. Using ABC allows us to identify governing parameters and mechanisms for root growth that are most supported by real observations.

We tested our ABC approach with synthetic observations from models of thale cress and narrowleaf lupin, confirming that we can recover the parameter values we put in. We then used real thale cress plants (wild and mutant) to show that our platform distinguishes genetically different plants and identifies most-likely parameters and model structures for real root growth. We used the platform to select models for growth and branching, showing how it can be used to compare existing models from the literature. We hope that this approach can be used to further help improve the interpretability and rigour of plant modelling and simulation! Iain and Clare

Saturday, 22 September 2018

ARTICLE: Time marches on -- mitochondria, ageing, and disease

Burgstaller, J.P., Kolbe, T., Havlicek, V., Hembach, S., Poulton, J., Piálek, J., Steinborn, R., Rülicke, T., Brem, G., Jones, N.S. and Johnston, I.G. Large-scale genetic analysis reveals mammalian mtDNA heteroplasmy dynamics and variance increase through lifetimes and generations. Nature communications2488 (2018)

DNA in mitochondria, the powerhouses of the cell, is passed down from mother to child. But there are many mitochondria in each cell, and these mitochondria may have different genetic features. If a mother carries a mixture of mitochondrial DNA (mtDNA) types, this can make it hard to say which features their children will inherit. For mothers carrying a disease-causing mtDNA mutation, this makes family planning and clinical therapies challenging.


In particular, the role of a mother's age has long been a mystery. Is the probability of a child inheriting a particular mtDNA feature higher when mothers are younger or older? An answer to this question could help plan clinical strategies to improve fertility and prevent the inheritance of deadly mitochondrial disease.


To address this, we worked with our excellent collaborators with a combination of maths, statistics, and experiment. Our collaborators used cutting-edge technology to reveal the mixtures of mtDNA in the egg cells of mother mice at a wide range of ages, and in the litters of offspring the mothers produced. This experimental work was the largest-scale study of mammalian mtDNA that we're aware of, involving thousands of observations throughout lifetimes and between generations. In concert, we developed a mathematical model describing the changes to, and inheritance of, mtDNA from mother to offspring. We combined the model and data to learn how different biological processes affect mtDNA through and between generations.



Cells contain populations of mitochondria, and these populations change over time. In European mice, we observed how variability in these populations evolves as mammals age and reproduce. We found that older mother have more varied mitochondria and pass this variance on to their offspring -- of central importance in the inheritance of genetic disease. 

We found that the variability of mtDNA dramatically increased as mothers aged. This means that the probability of inheriting more extreme -- both lower and higher -- levels of a genetic feature increases for older mothers. We also found that different mtDNA mixtures were inherited in different ways - with some mtDNA types favoured for inheritance and some disfavoured. We used our findings to create a way to predict how the risk that offspring would inherit disease-causing mtDNA features changes over time. Moving forward, we're aiming to harness these powerful ways of using large datasets to describe and predict the dynamics of mtDNA inheritance in humans, and to learn what it is about these mtDNA types that predicts their evolution across generations. You can read the article for free in Nature Communications here.

ARTICLE: How plants decide when to germinate

Topham, A.T., Taylor, R.E., Yan, D., Nambara, E., Johnston, I.G. and Bassel, G.W. Temperature variability is integrated by a spatially embedded decision-making center to break dormancy in Arabidopsis seeds. PNAS 114 6629 (2017)

A plant's choice to germinate is one of the most important decisions in the world. If it is made too soon, the plant may be damaged by harsh winter conditions; if too late, the plant may be outcompeted, and crop yields may be lower. If crops in a field make the decision at different times, there is more room for weeds to grow and pests to take over. 


In a recent study, we combined mathematical modelling with several neat experiments to identify sets of cells that make this germination choice in a much-studied plant called thale cress (Arabidopsis thaliana), and have learned how it makes decisions based on the plant's environment.



Two views of the plant embryo from laser microscopy, highlighting cells where different components of the germination control machinery are expressed. The background shows the "attractor basins" in a mathematical description of the germination decision: horizontal and vertical axes give the levels of two hormones ABA and GA, the blue region corresponds to dormant seeds and the red region to germination. 

This germination circuitry functions through a circuit of chemical stimuli and responses. Using laser microscopy, we found that different parts of this circuit exist in different parts of the plant embryo -- and that the separation of these parts is central to how the brain functions. We used mathematical modelling to show that communication between separated elements of the germination circuitry controls the plant's sensitivity to its environment. Following this theory, we used a mutant plant where cells were more chemically linked -- essentially enhancing communication between circuit elements -- to show that germination depends on these intra-cellular signals.


The separation of circuit elements allows a wider palette of responses to stimuli. It's like the difference between reading one critic's review of a film four times over, or amalgamating four different critics' views before deciding to go to the cinema. Our mathematical theory predicted that more plants would germinate when exposed to varying environments -- like three short pulses of cold -- than constant environments -- like one long cold period. We tested this theory in the lab and found exactly this behaviour.


Next, the hope is to learn about the germination brain in other plants and crops, and to show how our new knowledge of the germination machinery can be used to enhance and synchronise germination in crops. You can read the paper for free in the journal PNAS here. Iain

Saturday, 10 June 2017

ARTICLE: Supply, demand, energy, and death

Mitochondrial heterogeneity, metabolic scaling and cell death
J Aryaman, H Hoitzing, JP Burgstaller, IG Johnston, NS Jones
BioEssays e201700001; doi:10.1002/bies.201700001 (2017)
  •  The links between mitochondrial functionality and various aspects of cell physiology remain unclear; we combine recent experimental insights with mathematical modelling to produce quantitative hypotheses linking metabolism, cell proliferation, and mitochondria.
Cells need energy to produce functional machinery, deal with challenges, and continue to grow and divide -- these activities and others are collectively referred to as "cell physiology". Mitochondria are the dominant energy sources in most of our cells, so we'd expect a strong link between how well mitochondria perform and cell physiology. Indeed, when mitochondrial energy production is compromised, deadly diseases can result -- as we've written about before.

The details of this link -- how cells with different mitochondrial populations may differ physiologically -- is not well understood. A recent article shed new light on this link by looking at a measure of mitochondrial functionality in cells of different sizes. They found what we'll call the "mitopeak" -- mitochondrial functionality peaks at intermediate cell sizes, with larger and smaller cells having less functional mitochondria. The subsequent interpretation was that there is an “optimal”, intermediate, size for cells. Above this size, it was suggested that a proposed universal relationship between the energy demands of organisms (from microorganisms to elephants) and their size predicts the reduction in the function of mitochondria. Smaller cells, which result from a large cell having divided, were suggested to have inherited their parent's low mitochondrial functionality. Cells were predicted to “reset” their mitochondrial activity as they initially grow and reach an “optimal” size.

We were interested in the mitopeak, and wondered if scientifically simpler hypotheses could account for it. Using mathematical modelling, our idea was to use the observation that as a cell becomes larger in volume, the size of its mitochondrial population (and hence power supply) increases in concert. We considered that a cell has power demands which also track its volume, as well as demands which are proportional to surface area and power demands which do not depend on cell size at all (such as the energetic cost of replicating the genome at cell division, since the size of a cell's genome does not depend on how big the cell is). Assuming that power supply = demand in a cell, then bigger cells may more easily satisfy e.g. the constant power demands. This is because the number of mitochondria increases with cell volume yet the constant demands remain the same regardless of cell size. In other words, if a cell has more mitochondria as it gets larger, then each mitochondrion has to work less hard to satisfy power demand.

To explain why the smallest cells also have mitochondria which do not appear to work hard, we suggested that some smaller cells could be in the process of dying. If smaller cells are more likely to die, and if dying cells have low mitochondrial functionality (both of these ideas are biologically supported), then, by combining this with the power supply/demand picture above, the observed mitopeak naturally emerges from our mathematical model.

As an alternative model, we also suggested that the mitopeak could come entirely from a nonlinear relationship between cell size and cell death, with mitochondrial functionality as a passive indicator of how healthy a cell is. This indicates the existence of multiple hypotheses which could explain this new dataset.


A recent study has provided new data for the relationship between cell physiology and mitochondrial functionality. We have used mathematical modelling to suggest that a mixture of cellular power demand scaling, as well as cell death, could intuitively account for these new data. However, a nonlinear relationship between cell death and cell size could also account for these data, as well as a nonlinear relationship between mitochondrial functionality and cell size, as proposed by the original authors of the dataset. By integrating such a relationship between cell size and mitochondrial functionality into one of our existing models, we found that this “mitopeak” helps explain a wider set of cell physiological data. Using our model to highlight these competing hypotheses, we suggest future experiments to gather further support for these potential explanations.

Interestingly, we also found that the mitopeak could be an alternative to one aspect of a model we used some time ago to explain a different dataset, looking at the physiological influence of mitochondrial variability. Then, we modelled the activity of mitochondria as a quantity that is inherited identically by each daughter cell from its parent, plus some noise -- noting that this was a guess at the true behaviour because we didn't have the data to make a firm statement. We needed this relationship because observed functionality varied comparatively little between sister cells but substantially across a population. The mitopeak induces this variability without needing random inheritance of functionality, and may thus be the refined picture we've been looking for. These ideas, and suggestions for future strategies to explore the link between mitochondria and cell physiology in more detail, are in our new BioEssays article here. Juvid, Nick, and Iain.

Sunday, 21 May 2017

ARTICLE: A healthy dose of mathematics

Toward Precision Healthcare: Context and Mathematical Challenges
C Colijn, N Jones, IG Johnston, S Yaliraki, M Barahona
Frontiers in Physiology 8 136 (2017)
  • The continuing explosion of available biomedical data will help us tailor and optimise therapies for individual patients; we are designing new maths and statistics to help this process and to include social and other data into an overarching "precision healthcare" approach.
Our research combines tools from maths and statistics with biological data to learn more about the biological world. An exciting, growing, and much-discussed branch of science -- precision medicine -- is a specific instance of this idea. The vision of precision medicine is to use the expanding volume of data that's emerging from medicine and biology to tailor and optimise medical therapies for individual patients, making the therapies as effective as possible. This idea isn't new -- we are well aware, for example, that an individual's blood type dictates which blood transfusions they can successfully receive. But precision medicine is a much bigger picture, potentially taking into account large amounts of genetic, environmental, dietary, and other features to identify the optimal treatment for a disease -- for example, tailoring chemotherapy treatments to match the genetic specifics of a particular cancer case.

Dealing with these large and diverse datasets will need new mathematical and statistical approaches, built with an ongoing link to clinical practice. At the same time, we're interested in expanding the idea of precision medicine to include the "big data" that's increasingly available about individuals' social and logistic contexts. Social networks can dictate how diseases spread -- and how knowledge and views about therapies, vaccines, and other medically pertinent ideas are transmitted and shaped from person to person. A person's home region determines the genetic structure of local people who may act as donors. We're looking at the idea of "precision healthcare" -- using new maths and statistics to optimise healthcare strategy, not just individual therapies, in the light of large-scale datasets.





One aspect of precision healthcare we'll be exploring is exploring how progressive diseases -- those that involve the accumulation of symptoms over time -- develop in patients, using transition networks like those above to model "disease spaces" and find pathways in those spaces.

We're excited to be part of a new initiative -- the Centre for the Mathematics of Precision Healthcare -- involving six parallel and related research projects that align with this goal. Some of our previous work -- for example, estimating social structures of big UK cities to explore the challenges that genetic diversity poses to gene therapies for mtDNA disease -- already has a precision healthcare feel. In a new review paper (available for free) we discuss this and other examples of past and future work that we hope will contribute to the precision healthcare goal, along with some key ideas and context for the initiative. Iain

Monday, 31 October 2016

ARTICLE: The maths of mitochondrial DNA

Evolution of Cell-to-Cell Variability in Stochastic, Controlled, Heteroplasmic mtDNA Populations
IG Johnston, NS Jones
The American Journal of Human Genetics 99 (5), 1150-1162 (2016)
  • Vital populations of mtDNA are constantly evolving in our cells in response to random influences and control from the nucleus: we build a general mathematical theory describing this poorly-understood process and show that it predicts a wide range of existing experimental outcomes and gives us lots of new insights into biology and disease
Mitochondrial DNA (mtDNA) contains instructions for building important cellular machines. We have populations of mtDNA inside each of our cells -- almost like a population of animals in an ecosystem. Indeed, mitochondria were originally independent organisms, that billions of years ago were engulfed by our ancestor's cells and survived -- so the picture of mtDNA as a population of critters living inside our cells has evolutionary precedent! MtDNA molecules replicate and degrade in our cells in response to signals passed back and forth between mitochondria and the nucleus (the cell's "control tower"). Describing the behaviour of these population given the random, noisy environment of the cell, the fact that cells divide, and the complicated nuclear signals governing mtDNA populations, is challenging. At the same time, experiments looking in detail at mtDNA inside cells are difficult -- so predictive theoretical descriptions of these populations are highly valuable.

Why should we care about these cellular populations? MtDNA can become mutated, wrecking the instructions for building machines. If a high enough proportion of mtDNAs in a cell are mutated, our cells struggle and we get diseases. It only takes a few cells exceeding this "threshold" to cause problems -- so understanding the cell-to-cell distribution of mtDNA is medically important (as well as biologically fascinating). Simple mathematical approaches typically describe only average behaviours -- we need to describe the variability in mtDNA populations too. And for that, we need to account for the random effects that influence them.
 

In our cells, signals from the "control tower" nucleus lead to the replication (orange) and degradation (purple) of mtDNA. These processes affect mtDNA populations that may contain normal (blue) and mutant (red) molecules. Our mathematical approach -- extending work addressing a similar but simpler system -- describes how the total number of machines, and the proportion of mutants, is likely to behave and change with time and as cells divide.

In the past, we have used a branch of maths called stochastic processes to answer questions about the random behaviour of mtDNA populations. But these previous approaches cannot account for the "control tower" -- the nucleus' control of mtDNA. To address this, we've developed a mathematical tradeoff -- we make a particular assumption (which we show not to be unreasonable) and in exchange are able to derive a wealth of results about mtDNA behaviour under all sorts of different nuclear control signals. Technically, we use a rather magical-sounding tool called "Van Kampen's system size expansion" to approximate mtDNA behaviour, then explore how the resulting equations behave as time progresses and cells divide.

Our approach shows that the cell-to-cell variability in heteroplasmy (the potentially damaging proportion of mutants in a cell) generally increases with time, and surprisingly does so in the same way regardless of how the control tower signals the population. We're able to update a decades-old and commonly-used expression (often called the Wright formula) for describing heteroplasmy variance, so that the formula, instead of being rather abstract and hard to interpret, is directly linked to real biological quantities. We also show that control tower attempts to decrease mutant mtDNA can induce more variability in the remaining "normal" mtDNA population. We link these and other results to biological applications, and show that our approach unifies and generalises many previous models and treatments of mtDNA -- providing a consistent and powerful theoretical platform with which to understand cellular mtDNA populations. The article is in the American Journal of Human Genetics here and a preprint version can be viewed here. Iain

Thursday, 18 February 2016

ARTICLE: Who keeps the plans for our power stations?

Evolutionary Inference across Eukaryotes Identifies Specific Pressures Favoring Mitochondrial Gene Retention

IG Johnston, BP Williams
Cell Systems 2 (2), 101-111 (2016)

  • Why some genes are retained in mitochondria, where they are prone to disease-causing mutation, is a much-debated evolutionary question: we use new and generalisable maths and statistics to harness a large volume of sequence data and find the features of genes that predict the patterns of mitochondrial evolution that we observe.


Billions of years ago, a single-celled organism that would become our ancestor engulfed another smaller single-celled organism. The engulfed cell was probably intended to be lunch, but for reasons that remain mysterious (though recently explored here), it remained intact within our ancestor. It produced valuable chemicals that our ancestor could make use of, and was protected within the larger cell. This started a mutually beneficial relationship that evolved over billions of years to give rise to our situation today -- we are the descendants of the big cell, and our mitochondria are the descendants of the small, engulfed cell. 

As they were once independent organisms, mitochondria possess their own genomes (mitochondrial DNA, or mtDNA, which we've written about before). However, unlike the genomes of independent single-celled organisms like bacteria, mtDNA has only a handful of genes: why? Over evolutionary time, the majority of genes have either vanished from mtDNA or been transferred to the nucleus of the host cell. The reasons for transferring these genes to the nucleus are quite well understood; the nucleus is a safer environment for genes, less prone to mutation, and has several other evolutionary advantages.  But, given that transfer to the nucleus is possible, and genes in mtDNA are susceptible to mutation and damage (often giving rise to devastating diseases, which we study and try to prevent), why have mitochondria retained any genes at all? 

This question has been asked for decades, but until recently we lacked the data and the mathematical language to answer it quantitatively. Scientists energetically debate several different hypotheses: our approach attempts to let the data speak for itself without any preconceived ideas about which hypotheses are most likely. To this end, we built a mathematical model encompassing the evolutionary history of organisms with mitochondria, and a powerful statistical framework to amalgamate all the data that has been collected in recent years -- thousands of mitochondrial genomes from organisms from plants to protists (and humans) -- and harness it to compare the many disputed hypotheses addressing this question. 

Our mathematical approach allows us to "rewind the tape of evolution" and explore how mitochondrial genes have evolved. We're looking at Complex I -- an important protein complex involved in respiration -- over time, and watching the number of its subunits encoded in mitochondrial DNA (coloured black) decrease over evolutionary time, according to rules which we identify. The skyscrapers in the background are part of a graph describing how more mtDNA genes are lost over evolutionary history.

In a new paper in Cell Systems here (free here) we found several features that are most related to whether a gene is retained in mtDNA. Before discussing what they were, note that this picture -- several different features each with some influence -- explains and justifies the existing scientific debate. If hypothesis X and hypothesis Y both represent parts of the underlying "truth", then scientists advocating X alone and scientists advocating Y alone are neither completely wrong nor necessarily at odds -- everyone's partly right and the truth lies in the combination of the two arguments. 

The features that predict mtDNA gene retention are how central a gene's product is in its protein complex, the hydrophobicity of the protein the gene encodes, and the proportion of G's and C's in the gene's sequence. This suggests that genes are retained in mtDNA:  
(a) To allow local control of mitochondrial machinery (individual mitochondria can be controlled in response to cellular demands, rather than having to apply changes to the entire cellular population of mitochondria at once).  
(b) To prevent hydrophobic proteins ending up in the wrong place in the cell (if encoded by the far-away nucleus, these proteins may not be able to reach or enter the mitochondrion). 
(c and most speculatively) Because they are capable of withstanding the damaging environment of the mitochondria (GC-rich DNA and RNA is chemically more robust than GC-poor molecules). 

We found that the combination of the features we identified also predicted the success of experiments where scientists have attempted to mimic evolution and artificially transfer genes from the mitochondrion to the nucleus. Our results, as well as addressing a central mystery of evolutionary biology, thus also have the potential to inform synthetic biology approaches to tailor the genetics and bioenergetics of organisms. One final but important point is that the mathematical and statistical machinery we built for this project is highly generalisable and an efficient way of harnessing large sets of data about evolutionary and progressive processes -- we hope to use it to explore lots of other questions, including figuring out the pathways of disease progression and suggesting personalised medicine strategies in the clinic. Iain and Ben

Wednesday, 27 January 2016

ARTICLE: How evolution deals with mitochondrial mutants (and how we can take advantage)

Stochastic modelling, Bayesian inference, and new in vivo measurements elucidate the debated mtDNA bottleneck mechanism

  • Disease-causing mutant mtDNA is inherited through a complicated process: we use maths and statistics to shed light on this process and suggest possible therapeutic strategies to address disease inheritance and onset
Our mitochondrial DNA (mtDNA) provides instructions for building vital machinery in our cells. MtDNA is inherited from our mothers, but the process of inheritance -- which is important in predicting and dealing with genetic disease -- is poorly understood. This is because mitochondrial behaviour during development (the process through which a fertilised egg becomes an independent organism) is rather complex. If a mother's egg cell begins with a mixed population of mtDNA -- say with some type A and some type B -- we usually observe hard-to-predict mtDNA differences between cells in the daughter. So if the mother's egg cell starts off with 20% type A, egg cells in the daughter could range (for example) from 10%-30% of type A, with each different cell having a different proportion of A. This increase in variability, referred to as the mtDNA bottleneck, is important for the inheritance of disease. It allows cells with higher proportions of mutant mtDNA to be removed; but also means that some cells in the next generation may contain a dangerous amount of mutant mtDNA. Crucially, how this increase in variability comes about during development is debated. Does variability increase because of random partitioning of mtDNAs at cell divisions? Is it due to the decreased number of mtDNAs per cell, increasing the magnitude of genetic drift? Or does something occur during later development to induce the variability? Without knowing this in detail, it is hard to propose therapies or make predictions addressing the inheritance of disease.

We set out to answer this question with maths! Several studies have provided data on this process by measuring the statistics of mixed mtDNA populations during development in mice. The different studies provided different interpretations of these results, proposing several different mechanisms for the bottleneck. We built a mathematical framework that was capable of modelling all the different mechanisms that had been proposed. We then used a statistical approach called approximate Bayesian computation to see which mechanism was most supported by the existing data. We identified a model where a combination of copy number reduction and random mtDNA duplications and deletions is responsible for the bottleneck. Exactly how much variability is due to each of these effects is flexible -- going some way towards explaining the existing debate in the literature.  We were also able to solve the equations describing the most likely model analytically. These solutions allow us to explore the behaviour of the bottleneck in detail, and we use this ability to propose several therapeutic approaches to increase the "power" of the bottleneck, and to increase the accuracy of sampling in IVF approaches.

A "bottleneck" acts to increase mtDNA variability between generations. But how is this bottleneck manifest? Our approach suggests that a combination of copy number reduction (pictured as a "true" copy number bottleneck), and later random turnover of mtDNA (pictured as replication and degradation), is responsible.

Our excellent experimental collaborators, lead by Joerg Burgstaller, then tested our theory by taking mtDNA measurements from a model mouse that differed from those used previously and which, could in principle have shown different behaviour. The behaviour they observed agreed very well with the predictions of our theory, providing encouraging validation that we have identified a likely mechanism for the bottleneck. New measurements also showed, interestingly, that the behaviour of the bottleneck looks similar in genetically diverse systems, providing evidence for its generality. You can read about this in the free (open-access) journal eLife here. Iain and Nick [blog article also here]

ARTICLE: Inferring the evolutionary history of photosynthesis : C 4 yourself


Phenotypic landscape inference reveals multiple evolutionary paths to C4 photosynthesis

  • Some plants have evolved efficient photosynthesis, but important crops including rice have not: we use maths and statistics in conjunction with biological data to understand this evolution, with a view to repeating it artificially in crops to increase food production

Biological evolution is a complex, stochastic process which dictates fundamental properties of life. Our understanding of evolutionary history is severely limited by the sparsity of the fossil record: we only have a handful of fossilised snapshots to infer how evolution may have progressed throughout the history of life. Many physicists and mathematicians have attempted theoretical treatments of the process of evolution, using varying degrees of abstraction, in order to provide a more solid quantitative foundation with which to study this complex and important phenomenon, but the predictive power of these theoretical models, and their ability to answer specific biological questions, is often questioned.

We recently focussed on one remarkable product of evolution in plants: so-called "C4 photosynthesis". C4 consists of a complex set of changes to the genetic and physiological features which have evolved in some plants and act to increase the efficiency of photosynthesis. This complex set of changes has evolved over 60 times convergently: that is, plants from many different lineages independently "discover" C4 photosynthesis through evolution. We were interested in the evolutionary history of how these discoveries occurred -- both motivated by fundamental biology and the possibility of "learning from evolution" and using information about the evolution of C4 to design more efficient crop plants.


C3 and C4 plants differ in several physical and genetic ways (leaves and cells either side). We picture evolution as progressing along paths over a hypercube connecting these states (grey lines) -- some paths will give rise to intermediate species matching those we really observe (red and blue points). We can calculate how likely each path is and thus reconstruct evolutionary history.
 
To this end, we modelled the evolution of C4 as a pathway through a space containing many different possible plant features. The pathway starts at C3 -- the precursor to C4 -- and progressively takes steps in different directions, acquiring one-by-one the features that sum up to C4 photosynthesis. Using a survey of plant properties from across the wide scientific literature, we identified which intermediate states these pathways were likely to pass through, given observed properties of plants that currently possess some, but not all, C4 features. We were then able to use a new inference technique to predict the ordering in which these likely pathways traverse the evolutionary space. We showed that this approach worked by both successfully inferring the known evolutionary steps in synthetic datasets and correctly predicting previously unknown properties of several plants, which we verified experimentally. Our (open access) paper is here and there's a less technical summary and commentary here. Our approach showed that C4 photosynthesis can evolve through a range of distinct evolutionary pathways, providing a potential explanation for its striking convergence. Several of these different pathways were made explicitly visible when we examined the inferred evolutionary histories of different plant lineages -- different families are likely to have converged on C4 through different evolutionary routes. Furthermore, the most likely initial steps towards C4 photosynthesis are surprisingly not directly related to photosynthesis, being solutions to different biological challenges, but also providing evolutionary "foundations" upon which the machinery of C4 can evolve further. We hope that the recipes for C4 photosynthesis that we have inferred find use in efficient crop design, and anticipate our inference procedure being of use in the study of other specific biological questions regarding evolutionary histories. Iain [blog article also here]