Showing posts with label noise. Show all posts
Showing posts with label noise. Show all posts

Wednesday, 8 January 2020

ARTICLE: Learning pathways of disease progression

HyperTraPS: Inferring probabilistic patterns of trait acquisition in evolutionary and disease progression pathways
Sam F Greenbury, Mauricio Barahona, Iain G Johnston
Cell Systems (2019)

Many diseases that take a substantial human toll can be viewed as “progressive”. That is, a patient starts out healthy, then disease-related problems and/or symptoms develop over time. For example, a given case of cancer may begin with a patient acquiring a particular mutation, then other mutations building up in their genome over time.

How the same disease progresses in different patients often varies widely. Understanding this variability is important for precision medicine, where detailed knowledge of individual patients is used to design the best targeted treatments. However, learning the varied pathways of diseases and using them to predict future outcomes is challenging. Human researchers usually cannot hope to remember or analyse enough examples of patient data to provide the most reliable picture.

We previously developed an algorithm called HyperTraPS (hypercubic transition path sampling) to explore how biological systems evolve over time. We reasoned that HyperTraPS could also be used to learn the pathways of disease progression. In a new study in Cell Systems (free preprint available here) we used HyperTraPS to analyse biomedical data from many patients – hundreds, or thousands of individuals – to build a ‘road map’ of the different pathways that a disease takes over time.

Picture a river that branches out into a wide delta. Patients start out healthy – upstream in the river – and different patients go down different branches as the disease progresses and they acquire more symptoms. HyperTraPS learns the structure of the river delta from data, and predicts which river branches are more or less likely – and, importantly, where you'll end up if you're currently at a particular point.

By learning these branching patterns of disease progression, HyperTraPS has helped provide a refined risk assessment for malaria, based on data from thousands of Gambian children – as we’ve written about before. The approach also revealed diverse pathways of ovarian cancer progression, where the first mutation to occur appears to play a large role in determining subsequent mutations.

The "waterfall" in the foreground shows paths from one stage of a disease to the next, learnt by HyperTraPS using data from a high number of patients. Each dot of the illustration represents different stages of disease, for example a specific set of symptoms or a given set of mutations. The thickness of the lines indicate the probability of moving from one specific stage of disease to the next.

HyperTraPS is very generalisable and can be used to learn pathways by which mutations, symptoms, or other features develop over time from an initial state. We further used this generalisability to understand a biomedically important example of evolution – specifically, how tuberculosis evolves to become resistant to antibiotics.

Tuberculosis acquires resistance through mutations, and HyperTraPS has revealed the patterns of these mutations in TB bacteria reported from a group of 1000 Russian patients. These patterns help predict which mutation a bacterium will acquire next, and hence which drugs may be more effective for a given case. We’re following up with other applications of HyperTraPS, to learn about other progressive diseases, ageing, and evolution, and even to analyse how students complete tasks in online courses.

Monday, 28 January 2019

ARTICLE: How mitochondria can vary, and consequences for human health

(cross-posted from Imperial Mitochondriacs)

Mitochondria are components of the cell which are involved in generating “energy currency” molecules called ATP across much of complex life. Since many mitochondria exist within single cells (often hundreds or thousands), it is possible for the characteristics of individual mitochondria to vary within cells, and within tissues. This variation of mitochondrial characteristics can affect biological function and human health.

Since mitochondria possess their own, small, circular, DNA molecules (mtDNA), we can split mitochondrial characteristics into two categories: genetic and non-genetic. In our review, we discuss a number of aspects in which mitochondria vary, from both genetic and non-genetic perspectives. 



In terms of mitochondrial genetics, the amount of mtDNA per cell is variable. When a cell divides, its daughters receive a share of its parents mtDNA, but the split isn’t precisely 50/50, so cell division can cause variability in the number of mtDNAs per cell. As mtDNAs are replicated and degraded over time, errors in the copying process may give rise to mtDNA mutations, which may spread throughout a cell. Factors such as: the total amount, the rate of degradation/replication, the mean fraction of mutants, and the extent of fragmentation in the mitochondrial network, can all influence how variable the fraction of mutated mtDNAs becomes through time (see here for a preview of some upcoming work on this topic). The total amount, and mutated fraction of mtDNAs, are implicated in diseases such as neurodegeneration, as well as the ageing process.

Apart from genetic variations, there are many non-genetic features of mitochondria which also vary within and between cells. Changes in mtDNA sequence can change the amino-acid sequence of the proteins encoded by mtDNA, causing structural changes in the molecular machines which generate ATP. The shape of the membranes of mitochondria are also highly variable, and respond to mitochondrial activity through quantities such as pH, where mitochondrial activity itself may depend on mtDNA sequence. The previous two examples (mitochondrial protein and membrane structure) demonstrate how the genetic state of mitochondria may influence their non-genetic characteristics. Mitochondrial non-genetic characteristics may also influence the genetic state: for instance, mitochondrial membrane potential can influence the probability of a mitochondria being degraded, along with its mtDNA.

The inter-dependence of genetic and non-genetic characteristics demonstrate the complex feedback loops linking these two aspects of mitochondrial physiology. We suggest here that, since changes in mitochondrial genetics occur more slowly than most physical aspects of mitochondrial physiology, understanding mitochondrial genetics may be especially important in explaining phenomena such as ageing, which appears to be closely related to mitochondrial heterogeneity. You can freely access our work, which has recently been published in Frontiers in Genetics, as “Mitochondrial Heterogeneity” https://www.frontiersin.org/articles/10.3389/fgene.2018.00718/full Juvid, Iain and Nick.
 

Saturday, 22 September 2018

ARTICLE: How do plants roll dice?

Johnston, I.G. and Bassel, G.W. Identification of a bet-hedging network motif generating noise in hormone concentrations and germination propensity in Arabidopsis. Journal of the Royal Society Interface15 141 (2018)

Seeds feed the world, and uniform, reliable harvests of seeds and grains is essential for food security. However, there's a fundamental tension between the evolutionary priorities of plants and the agricultural priorities of humans. Evolutionarily, it is good for plants to "hedge their bets" by having seeds germinate at different times. A plant whose seeds all germinate in March will be susceptible to a frost in April, potentially leading to the loss of a generation of offspring. By contrast, a plant whose seeds germinate throughout March and April will have a subset of its offspring survive that frost, and its genes will be passed on to the next generation.


This bet-hedging poses a challenge for agriculture. In agricultural settings, we have more control over plant environments, and so plants have less need to withstand unpredictable environmental fluctuations. At the same time, non-uniform germination decreases crop yields, makes harvesting harder, and makes crops more susceptible to pest invasion. If we can learn how plants generate this evolved germination variability, we can design engineering and/or breeding strategies to reduce this and improve crop yields.



Plants have evolved to "hedge their bets" by having seeds germinate at different times -- this makes generations of plants more robust to environmental fluctuations. Our work reveals a mechanism that "rolls dice" within plant cells, acting like a random number generator to produce variability in germination propensity. 

In a previous paper (blog post here), we looked at how germination is controlled by an interaction between two hormones known as ABA and GA. During that project, we noticed a surprising feature of the cellular pathways affecting ABA. Oddly, it seemed that ABA both activated a pathway that increased its own production, and at the same time (and in the same place) activated a pathways that increased its own degradation. These two pathways seemed to be competitive -- one increases levels of ABA, the other decreases them. Why would cells spend energy in this "futile" way?


We hypothesised that these competitive pathways might have the effect of generating variability in ABA levels. The pathways are fundamentally "noisy", involving random interactions in the chaotic environment of the cell. Consider increasing the activity of both pathways simultaneously. One pathway would act to increase levels of ABA, the other would act to decrease it. The increased "push and pull" of these noisy pathways would increase the spread of levels of ABA in different cells, even if average levels stayed the same.


Because it's hard to measure the levels of hormones in individual cells over time, we initially took a theoretical approach. We showed, with maths, that the competing pathways did indeed have this variability-inducing effect. By varying the activity through these pathways, the cell can increase variability in ABA levels, and hence increase variability in germination propensity. We showed that the theory we developed was compatible with some experiments where the ABA circuitry was artificially manipulated. The theory went on to reveal various aspects of cellular machinery that we could conceivably target through synthetic approaches, in order to reduce germination variability. Put together, our quantitative theory, supported by experiment, explained the mysterious competitive pathways and revealed several new interventions with the potential to improve food security. You can read about it for free in the Journal of the Royal Society Interface here. Iain  


ARTICLE: Which genes are essential for bacterial survival?

Goodall, E.C., Robinson, A., Johnston, I.G., Jabbari, S., Turner, K.A., Cunningham, A.F., Lund, P.A., Cole, J.A. and Henderson, I.R., 2018. The essential genome of Escherichia coli K-12. mBioe02096 (2018)

Bacteria cause diseases, and are developing resistance to the drugs we use to kill them. Anti-microbial resistance (AMR) is one of the most pressing global health challenges facing society. In the immense scientific endeavour of creating new, effective treatments for bacterial infections, fundamental biological knowledge about how bacteria live and proliferate is of vital importance.


One way we can obtain this knowledge is by discovering what cellular machinery that bacteria need to survive and proliferate. A common (and famous) bacterium called Escherichia coli (E. coli) has over 4000 protein-coding genes, but we're not really sure which of these genes is essential for the bacterium, and how many provide some non-essential "added value". If we can learn which genes are essential for bacteria, we have a more specific set of targets to shoot for in designing new drugs and therapies.


So -- how can we find out which genes are essential for E. coli? One neat way involves a new experimental approach called transposon-directed insertion site sequencing (TraDIS). Transposons are elements of DNA that can be inserted into a bacterial genome -- when they are inserted into part of the genome that codes for a gene, they prevent that gene being properly expressed, effectively removing it from the bacterium. TraDIS, in essence, takes a large population of bacteria and inserts one transposon into a random position in each bacterium. The population is then left to evolve for some time. After that time, we look at the genomes of bacteria within the surviving population, and see exactly where transposon insertions have been retained in some living bacteria.



A stylised representation of the E. coli genome and the positions within it where we found transposons to have been retained (corresponding to non-essential genes). 

The idea is that any bacteria in the population that have a transposon inserted into an essential gene will die. As such a gene is essential, it's required for survival, and a transposon preventing its expression will kill the bacterium. Therefore, if some bacteria in a population retain an insertion in gene X and survive, it follows that gene X is not essential. Conversely, if we see a large region of the genome within which no insertions are retained in the final population, it is likely that that region corresponds to an essential gene. 


There's some mathematical subtlety in the "it is likely". Depending on how many transposon insertions originally occur, and the length of the genome, some regions without insertions may occur just by chance. We did a bit of maths to work out how unlikely it is to see an insertion-free region of a given length arise by chance; and, by extension, how likely it is that a gene identified by this analysis is indeed essential for the bacterium. However, the maths was only one part of this project -- it was first and foremost an experimental tour de force by our excellent collaborators. We jointly provided a new atlas of essential genes in E. coli, provide a new way of reasoning about the powerful TraDIS technique, and provide several new insights into bacterial physiology and biochemistry. The work is freely available in the journal mBio here. Iain 

ARTICLE: How plants decide when to germinate

Topham, A.T., Taylor, R.E., Yan, D., Nambara, E., Johnston, I.G. and Bassel, G.W. Temperature variability is integrated by a spatially embedded decision-making center to break dormancy in Arabidopsis seeds. PNAS 114 6629 (2017)

A plant's choice to germinate is one of the most important decisions in the world. If it is made too soon, the plant may be damaged by harsh winter conditions; if too late, the plant may be outcompeted, and crop yields may be lower. If crops in a field make the decision at different times, there is more room for weeds to grow and pests to take over. 


In a recent study, we combined mathematical modelling with several neat experiments to identify sets of cells that make this germination choice in a much-studied plant called thale cress (Arabidopsis thaliana), and have learned how it makes decisions based on the plant's environment.



Two views of the plant embryo from laser microscopy, highlighting cells where different components of the germination control machinery are expressed. The background shows the "attractor basins" in a mathematical description of the germination decision: horizontal and vertical axes give the levels of two hormones ABA and GA, the blue region corresponds to dormant seeds and the red region to germination. 

This germination circuitry functions through a circuit of chemical stimuli and responses. Using laser microscopy, we found that different parts of this circuit exist in different parts of the plant embryo -- and that the separation of these parts is central to how the brain functions. We used mathematical modelling to show that communication between separated elements of the germination circuitry controls the plant's sensitivity to its environment. Following this theory, we used a mutant plant where cells were more chemically linked -- essentially enhancing communication between circuit elements -- to show that germination depends on these intra-cellular signals.


The separation of circuit elements allows a wider palette of responses to stimuli. It's like the difference between reading one critic's review of a film four times over, or amalgamating four different critics' views before deciding to go to the cinema. Our mathematical theory predicted that more plants would germinate when exposed to varying environments -- like three short pulses of cold -- than constant environments -- like one long cold period. We tested this theory in the lab and found exactly this behaviour.


Next, the hope is to learn about the germination brain in other plants and crops, and to show how our new knowledge of the germination machinery can be used to enhance and synchronise germination in crops. You can read the paper for free in the journal PNAS here. Iain

Saturday, 10 June 2017

ARTICLE: Supply, demand, energy, and death

Mitochondrial heterogeneity, metabolic scaling and cell death
J Aryaman, H Hoitzing, JP Burgstaller, IG Johnston, NS Jones
BioEssays e201700001; doi:10.1002/bies.201700001 (2017)
  •  The links between mitochondrial functionality and various aspects of cell physiology remain unclear; we combine recent experimental insights with mathematical modelling to produce quantitative hypotheses linking metabolism, cell proliferation, and mitochondria.
Cells need energy to produce functional machinery, deal with challenges, and continue to grow and divide -- these activities and others are collectively referred to as "cell physiology". Mitochondria are the dominant energy sources in most of our cells, so we'd expect a strong link between how well mitochondria perform and cell physiology. Indeed, when mitochondrial energy production is compromised, deadly diseases can result -- as we've written about before.

The details of this link -- how cells with different mitochondrial populations may differ physiologically -- is not well understood. A recent article shed new light on this link by looking at a measure of mitochondrial functionality in cells of different sizes. They found what we'll call the "mitopeak" -- mitochondrial functionality peaks at intermediate cell sizes, with larger and smaller cells having less functional mitochondria. The subsequent interpretation was that there is an “optimal”, intermediate, size for cells. Above this size, it was suggested that a proposed universal relationship between the energy demands of organisms (from microorganisms to elephants) and their size predicts the reduction in the function of mitochondria. Smaller cells, which result from a large cell having divided, were suggested to have inherited their parent's low mitochondrial functionality. Cells were predicted to “reset” their mitochondrial activity as they initially grow and reach an “optimal” size.

We were interested in the mitopeak, and wondered if scientifically simpler hypotheses could account for it. Using mathematical modelling, our idea was to use the observation that as a cell becomes larger in volume, the size of its mitochondrial population (and hence power supply) increases in concert. We considered that a cell has power demands which also track its volume, as well as demands which are proportional to surface area and power demands which do not depend on cell size at all (such as the energetic cost of replicating the genome at cell division, since the size of a cell's genome does not depend on how big the cell is). Assuming that power supply = demand in a cell, then bigger cells may more easily satisfy e.g. the constant power demands. This is because the number of mitochondria increases with cell volume yet the constant demands remain the same regardless of cell size. In other words, if a cell has more mitochondria as it gets larger, then each mitochondrion has to work less hard to satisfy power demand.

To explain why the smallest cells also have mitochondria which do not appear to work hard, we suggested that some smaller cells could be in the process of dying. If smaller cells are more likely to die, and if dying cells have low mitochondrial functionality (both of these ideas are biologically supported), then, by combining this with the power supply/demand picture above, the observed mitopeak naturally emerges from our mathematical model.

As an alternative model, we also suggested that the mitopeak could come entirely from a nonlinear relationship between cell size and cell death, with mitochondrial functionality as a passive indicator of how healthy a cell is. This indicates the existence of multiple hypotheses which could explain this new dataset.


A recent study has provided new data for the relationship between cell physiology and mitochondrial functionality. We have used mathematical modelling to suggest that a mixture of cellular power demand scaling, as well as cell death, could intuitively account for these new data. However, a nonlinear relationship between cell death and cell size could also account for these data, as well as a nonlinear relationship between mitochondrial functionality and cell size, as proposed by the original authors of the dataset. By integrating such a relationship between cell size and mitochondrial functionality into one of our existing models, we found that this “mitopeak” helps explain a wider set of cell physiological data. Using our model to highlight these competing hypotheses, we suggest future experiments to gather further support for these potential explanations.

Interestingly, we also found that the mitopeak could be an alternative to one aspect of a model we used some time ago to explain a different dataset, looking at the physiological influence of mitochondrial variability. Then, we modelled the activity of mitochondria as a quantity that is inherited identically by each daughter cell from its parent, plus some noise -- noting that this was a guess at the true behaviour because we didn't have the data to make a firm statement. We needed this relationship because observed functionality varied comparatively little between sister cells but substantially across a population. The mitopeak induces this variability without needing random inheritance of functionality, and may thus be the refined picture we've been looking for. These ideas, and suggestions for future strategies to explore the link between mitochondria and cell physiology in more detail, are in our new BioEssays article here. Juvid, Nick, and Iain.

Monday, 31 October 2016

ARTICLE: The maths of mitochondrial DNA

Evolution of Cell-to-Cell Variability in Stochastic, Controlled, Heteroplasmic mtDNA Populations
IG Johnston, NS Jones
The American Journal of Human Genetics 99 (5), 1150-1162 (2016)
  • Vital populations of mtDNA are constantly evolving in our cells in response to random influences and control from the nucleus: we build a general mathematical theory describing this poorly-understood process and show that it predicts a wide range of existing experimental outcomes and gives us lots of new insights into biology and disease
Mitochondrial DNA (mtDNA) contains instructions for building important cellular machines. We have populations of mtDNA inside each of our cells -- almost like a population of animals in an ecosystem. Indeed, mitochondria were originally independent organisms, that billions of years ago were engulfed by our ancestor's cells and survived -- so the picture of mtDNA as a population of critters living inside our cells has evolutionary precedent! MtDNA molecules replicate and degrade in our cells in response to signals passed back and forth between mitochondria and the nucleus (the cell's "control tower"). Describing the behaviour of these population given the random, noisy environment of the cell, the fact that cells divide, and the complicated nuclear signals governing mtDNA populations, is challenging. At the same time, experiments looking in detail at mtDNA inside cells are difficult -- so predictive theoretical descriptions of these populations are highly valuable.

Why should we care about these cellular populations? MtDNA can become mutated, wrecking the instructions for building machines. If a high enough proportion of mtDNAs in a cell are mutated, our cells struggle and we get diseases. It only takes a few cells exceeding this "threshold" to cause problems -- so understanding the cell-to-cell distribution of mtDNA is medically important (as well as biologically fascinating). Simple mathematical approaches typically describe only average behaviours -- we need to describe the variability in mtDNA populations too. And for that, we need to account for the random effects that influence them.
 

In our cells, signals from the "control tower" nucleus lead to the replication (orange) and degradation (purple) of mtDNA. These processes affect mtDNA populations that may contain normal (blue) and mutant (red) molecules. Our mathematical approach -- extending work addressing a similar but simpler system -- describes how the total number of machines, and the proportion of mutants, is likely to behave and change with time and as cells divide.

In the past, we have used a branch of maths called stochastic processes to answer questions about the random behaviour of mtDNA populations. But these previous approaches cannot account for the "control tower" -- the nucleus' control of mtDNA. To address this, we've developed a mathematical tradeoff -- we make a particular assumption (which we show not to be unreasonable) and in exchange are able to derive a wealth of results about mtDNA behaviour under all sorts of different nuclear control signals. Technically, we use a rather magical-sounding tool called "Van Kampen's system size expansion" to approximate mtDNA behaviour, then explore how the resulting equations behave as time progresses and cells divide.

Our approach shows that the cell-to-cell variability in heteroplasmy (the potentially damaging proportion of mutants in a cell) generally increases with time, and surprisingly does so in the same way regardless of how the control tower signals the population. We're able to update a decades-old and commonly-used expression (often called the Wright formula) for describing heteroplasmy variance, so that the formula, instead of being rather abstract and hard to interpret, is directly linked to real biological quantities. We also show that control tower attempts to decrease mutant mtDNA can induce more variability in the remaining "normal" mtDNA population. We link these and other results to biological applications, and show that our approach unifies and generalises many previous models and treatments of mtDNA -- providing a consistent and powerful theoretical platform with which to understand cellular mtDNA populations. The article is in the American Journal of Human Genetics here and a preprint version can be viewed here. Iain

Friday, 28 October 2016

ARTICLE: Random number seed

Variability in seeds: biological, ecological, and agricultural implications 
J Mitchell, IG Johnston, GW Bassel
Journal of Experimental Botany, erw397 (2016) 
  • Natural variability across scales, from the molecular to the environmental, means that individual seeds behave differently; we explore the challenges this variability poses for agriculture and food security, and how modern science can help address these challenges.
Seeds feed the world. Whether eaten themselves, or allowed to develop into crop plants which are then consumed by humans or livestock, seeds are the fundamental starting point for agriculture. But each seed has a different story. Throughout millions of years of evolution, plants have evolved to -- forgive the pun -- "hedge" their bets from one generation to the next. A parent plant cannot completely predict the environmental conditions that its offspring will face, so it induces variability in the seeds it produces. If some seeds are better at surviving in environment A and some are better in environment B, the plant has a way of ensuring its genes will survive regardless of whether the environment is A-like or B-like in future.

This bet-hedging is a sensible evolutionary strategy when environments are unpredictable. But modern agriculture makes environments much more predictable than the wild situations plants have been exposed to throughout evolutionary history. Now bet-hedging becomes a bad thing -- if we know the environment will always be C, energy spent ensuring that seeds survive in environments A and B is wasted, reducing potential yields.

Understanding and controlling the variability within populations of seeds thus has huge implications for agriculture. Variability inherent within populations of seeds, in addition to differences in the environments that seeds experience, means that, for example, seed lots germinate asynchronously (some quickly, some slowly or not at all). This leads to non-uniform and sub-optimal crop production, allows pests to enter fields, and challenges our ability to plan agricultural strategies. If we could control seed variability, these problems would be diminished, with a host of positive consequences for food security.

A given set of seeds will vary in their behaviour due to influences on many scales, from random molecular processes within cells to large-scale environmental stimuli. As a result, important features like germination propensity vary across seed lots (perhaps taking a broad distribution like that illustrated here), posing a challenge to agriculture and food security, which scientific understanding can mitigate.

In a new review, we survey our current understanding of the sources of variability in seeds, and its biological and agricultural implications. Processes across many scales induce variability in seed behaviour, from random cell biological interactions (like we've written about before!), through seed position in a parent plant, to large-scale environmental differences. We particularly focus on germination, an aspect of seed behaviour of crucial biological and agronomic importance, which takes place when a "developmental switch" in a seed is flipped. We discuss the genetic and molecular players that modern science has discovered to influence this decision to germinate in seeds, and describe the challenges in furthering our understanding of this vital question -- and how cool new tech, and maths, can help us make new progress! The review is in the Journal of Experimental Botany here. Iain

Wednesday, 31 August 2016

ARTICLE: Understanding the strength and correlates of immunisation programmes

Forecasted trends in vaccination coverage and correlations with socioeconomic factors: a global time-series analysis over 30 years

  •  Lack of trust in vaccines results in preventable illness and death all over the world; we use tools from statistics and large-scale socio-economic data to explore which features of a country "prime" it for weakened vaccine coverage, identifying factors which may help policymakers address vaccine confidence issues.
Childhood vaccinations are vital for the protection of children against dreadful diseases such as measles, polio, and diphtheria. In addition to providing personal protection, vaccines can also suppress epidemic outbreaks if a sufficiently large proportion of the population has immunity status – this “herd immunity” is important for society as many individuals are unable to vaccinate for medical reasons. Over the past half a century, public health organisations have made concerted efforts to vaccinate every child worldwide. However, notwithstanding the substantial improvements to vaccine coverage rates across the globe over the past few decades, there are still millions of unvaccinated children worldwide. The majority of these children live in countries where large numbers of the populations live in deprived, rural regions with poor access to healthcare. However, a number of children are denied vaccines because of parental attitudes and beliefs (which are often influenced by the media, religious groups, or anti-vaccination groups) – such hesitancy has been responsible for recent outbreaks in developing (e.g. Nigeria, Pakistan, Afghanistan) and developed (e.g. USA, UK) countries alike. Monitoring vaccine coverage rates, summarising recent vaccination behaviours, and understanding the factors which drive vaccination behaviour are thus key to our understanding vaccine acceptance, and can allow immunisation programmes to be more effectively tailored.

To understand these pertinent issues, we used machine learning tools on publicly-available vaccination and socioeconomic data (which can be found here and on the World Health Organization’s websites). We used Gaussian process regression to forecast vaccine coverage rates and used the predictive distributions over forecasted coverage rates to introduce a quantitative marker summarising a country’s recent vaccination trends and variability:  this summary is termed the Vaccine Performance Index. Parameterisations of this index can then be used to identify countries which are likely (over next few years) to have vaccine coverage rates far from those required for herd immunity levels or that are displaying worrying declines in rates and to assess which countries will miss immunisation goals set by global public health bodies. We find that these poorly-performing countries were mostly located in South-East Asia and sub-Saharan Africa though, surprisingly, a handful of European countries also perform poorly.




To investigate the factors associated with vaccination coverage, we sought links between socioeconomic factors with vaccine coverage and found that countries with higher levels of births attended by skilled health staff, gross domestic product, government health spending, and higher education levels have higher vaccination coverage levels (though these results are region-dependent).

Our vaccine performance index could aid policy makers’ assessments of the strength and resilience of immunisation programmes. Further,  identification of socioeconomic correlates of vaccine coverage points to factors to address to improve vaccination coverage. You can read further in our freely available paper – which is in collaboration with the London School of Hygiene and Tropical Medicine (Heidi Larson and David Smith) and IIT Delhi (Sumeet Agarwal) – in the open-access journal Lancet Global Health under the title “Forecasted trends in vaccination coverage and correlations with socioeconomic factors: a global time-series analysis over 30 years” and there is another free article unpacking it under the title "Global Trends in Vaccination Coverage". Alex, Iain, Nick.

Wednesday, 27 January 2016

ARTICLE: Generations of generating functions in dividing cells

Closed-form stochastic solutions for non-equilibrium dynamics and inheritance of cellular components over many cell divisions

  • Populations of important machines in our cells behave quite randomly: we build a mathematical framework to better understand these populations (which has already helped us understand the inheritance of mtDNA disease)
Cell biology is a unpredictable world, as we've written about before. The important machines in our cells replicate and degrade in processes that can be described as random; and when cells divide, the partitioning of these machines between the resulting cells also looks random. The number of machines we have in our cells is important, but how can we work with numbers in this unpredictable environment?

In our cells, machines are produced (red), replicate (orange), and degrade (purple) randomly with time, as well as being randomly partitioned when cells split and divide (blue). Our mathematical approach describes how the total number of machines is likely to behave and change with time and as cells divide.

Tools called "generating functions" are useful in this situation. A generating function is a mathematical function (like G(z) = z2, but generally more complicated) that encodes all the information about a random system. To find the generating function for a particular system, one needs to consider all the random things that can happen to change the state of that system, write them down in an equation (the "master equation") describing them all together, then use a mathematical trick to push that equation into a different mathematical space, where it is easier to solve. If that "transformed" equation can be solved, the result is the generating function, from which we can then get all the information we could want about a random system: the behaviour of its mean and variance, the probability of making any observation at any time, and so on.

We've gone through this mathematical process for a set of systems where individual cellular machines can be produced, replicated, and degraded randomly, and split at cell divisions in a variety of different ways. The generating functions we obtain allow us to follow this random cellular behaviour in new detail. We can make probabilistic statements about any aspect of the system at any time and after any number of cell divisions, instead of relying on assumptions that the system has somehow reached an equilibrium, or restricting ourselves to a single or small number of divisions. We've applied this tool to questions about the random dynamics of mitochondrial DNA (which we're very interested in! And this work connects explicitly with our recent eLife paper) in cells that divide (like our cells) or "bud" (like yeast cells), but the approach is very general and we hope it will allow progress in many more biological situations. You can read about this, free, here in the Proceedings of the Royal Society A. Iain and Nick [blog article also here]

ARTICLE: How evolution deals with mitochondrial mutants (and how we can take advantage)

Stochastic modelling, Bayesian inference, and new in vivo measurements elucidate the debated mtDNA bottleneck mechanism

  • Disease-causing mutant mtDNA is inherited through a complicated process: we use maths and statistics to shed light on this process and suggest possible therapeutic strategies to address disease inheritance and onset
Our mitochondrial DNA (mtDNA) provides instructions for building vital machinery in our cells. MtDNA is inherited from our mothers, but the process of inheritance -- which is important in predicting and dealing with genetic disease -- is poorly understood. This is because mitochondrial behaviour during development (the process through which a fertilised egg becomes an independent organism) is rather complex. If a mother's egg cell begins with a mixed population of mtDNA -- say with some type A and some type B -- we usually observe hard-to-predict mtDNA differences between cells in the daughter. So if the mother's egg cell starts off with 20% type A, egg cells in the daughter could range (for example) from 10%-30% of type A, with each different cell having a different proportion of A. This increase in variability, referred to as the mtDNA bottleneck, is important for the inheritance of disease. It allows cells with higher proportions of mutant mtDNA to be removed; but also means that some cells in the next generation may contain a dangerous amount of mutant mtDNA. Crucially, how this increase in variability comes about during development is debated. Does variability increase because of random partitioning of mtDNAs at cell divisions? Is it due to the decreased number of mtDNAs per cell, increasing the magnitude of genetic drift? Or does something occur during later development to induce the variability? Without knowing this in detail, it is hard to propose therapies or make predictions addressing the inheritance of disease.

We set out to answer this question with maths! Several studies have provided data on this process by measuring the statistics of mixed mtDNA populations during development in mice. The different studies provided different interpretations of these results, proposing several different mechanisms for the bottleneck. We built a mathematical framework that was capable of modelling all the different mechanisms that had been proposed. We then used a statistical approach called approximate Bayesian computation to see which mechanism was most supported by the existing data. We identified a model where a combination of copy number reduction and random mtDNA duplications and deletions is responsible for the bottleneck. Exactly how much variability is due to each of these effects is flexible -- going some way towards explaining the existing debate in the literature.  We were also able to solve the equations describing the most likely model analytically. These solutions allow us to explore the behaviour of the bottleneck in detail, and we use this ability to propose several therapeutic approaches to increase the "power" of the bottleneck, and to increase the accuracy of sampling in IVF approaches.

A "bottleneck" acts to increase mtDNA variability between generations. But how is this bottleneck manifest? Our approach suggests that a combination of copy number reduction (pictured as a "true" copy number bottleneck), and later random turnover of mtDNA (pictured as replication and degradation), is responsible.

Our excellent experimental collaborators, lead by Joerg Burgstaller, then tested our theory by taking mtDNA measurements from a model mouse that differed from those used previously and which, could in principle have shown different behaviour. The behaviour they observed agreed very well with the predictions of our theory, providing encouraging validation that we have identified a likely mechanism for the bottleneck. New measurements also showed, interestingly, that the behaviour of the bottleneck looks similar in genetically diverse systems, providing evidence for its generality. You can read about this in the free (open-access) journal eLife here. Iain and Nick [blog article also here]

ARTICLE: Great technological power, great statistical responsibility

Multiple hypothesis correction is vital and undermines reported mtDNA links to diseases including AIDS, cancer, and Huntingdon’s

  • Several papers perform incorrect and misleading statistical analyses in seeking links between mtDNA and cancer: these statistical issues must be corrected before scientific and policy progress can be made from these investigations
Biologists often report a result as a "significant" sign of exciting new science if there is less than a 1-in-20 chance that the result they observe could have emerged by chance from boring old science. This is silly (although we do it too!) -- by contrast, for example, physicists require less than a 1-in-3,500,000 chance. But this post won't discuss too many problems with this state of affairs -- that is done admirably elsewhere.

The problem can be compounded when scientists take lots of measurements. Say we take 50 measurements of a boring old system, and every time we see something that has less than a 1-in-20 chance of appearing in a boring old system, we call it "significant". We're playing the odds 50 times, so we expect to see 1-in-20 results appear around 2 or 3 times; just as if we roll a dice 50 times, we'd expect to roll a good few sixes. If we call every 1-in-20 result "significant" without accounting for the fact that we've looked at lots of measurements (and are thus more likely to see 1-in-20s by chance), we are in danger of reporting exciting new science when in fact the boring old science has been true all along.

There are lots of ways of doing this accounting, but a series of papers that have been recently published linking mtDNA to diseases have made no attempt to do it. Generally, these papers look at the mtDNA of people without the disease and the mtDNA of people with the disease. If any mtDNA features appear more in the people with the disease, the paper calculates the chance of that difference occurring in the boring old picture (in which there is no link between the mtDNA feature and the disease). If they drop below the 1-in-20 mark, they report an exciting new link between that feature and the disease. But they test dozens of features and never account for this multiple testing -- so, as above, we'd expect them to see "significant" results emerging just by chance. In a paper in Mitochondrial DNA here (free here) I show, by creating artificial data, that this problem is rife, that most of these reported links are spurious, and that scientists really need to be more responsible, before their flawed analysis starts to misguide health policy and medicine.

The top graph shows how the probability of seeing a 1-in-20 occurrence (p < 0.05 in the jargon), when in fact there is nothing new and exciting to report, increases as a scientist investigates more things. If an experiment consists of one test, then a 1-in-20 occurrence indeed has a 1-in-20 probability (0.05). But as soon as we do more tests, the chance of seeing at least one 1-in-20 occurrence starts to increase, as we are "playing the game" more times. If we do 6 tests there is a 0.27 probability -- between a 1-in-4 and 1-in-3 chance -- that we will see at least one 1-in-20 event. This is illustrated below, where we have six dice and think some of them may be unfair. We roll each one five times and count the number of 6s. One of them comes up 6 three times -- the chance of this happening for one fair die is less than 1-in-20. But because we've looked at six dice, we should be less surprised to see this rare event, because we've looked at more events in total. We need more evidence to claim that this die is unfair.

This quick note only represents the tip of the iceberg. MtDNA studies are often statistically unsound; statistical misdemeanours in biomedical studies are so common that most published research is wrong; scientists increasingly focus on the 1-in-20 chance as opposed to the size and importance of the effect they're measuring; the majority of hallmark papers in vital fields like cancer science are unreproducible (though this last point may have other causes than statistical problems). The 1-in-20 idea was only ever meant to be a step in identifying interesting scientific avenues, not the final measure of scientific truth. This is a big, and growing, problem! Iain

(For accessibility I have used "exciting", "boring", and "1-in-20" instead of their usual, more technical labels; they of course are usually called the "alternative hypothesis", "null hypothesis", and "p < 0.05" respectively).

ARTICLE: Turbocharging the back of the envelope

Explicit tracking of uncertainty increases the power of quantitative rule-of-thumb reasoning in cell biology

  • Estimated numbers in biology (and life) are often uncertain: we've made a calculator to work with this uncertainty and help make calculations more interpretable (with a particular focus on understanding how the cell works)
The numbers that we use to describe the world are rarely exact. How long will it take you to drive to work? Perhaps "between 20 and 30 minutes". It would be unwise (and unnecessary) to say "exactly 23.4 minutes".

This uncertainty means that "back-of-the-envelope" calculations are very valuable in estimating and reasoning about numerical problems, particularly in the sciences. The idea here is to perform a calculation using rough guesses of the quantities involved, to get an "order of magnitude" estimate of the answer you're after. Made famous in physics as "Fermi problems", attributed to Enrico Fermi (who used rough reasoning to deduce quantities from the power of an atomic bomb to the number of piano tuners in Chicago), this approach is integral in many current applications of maths and science. Cool books like "Street-fighting Mathematics", "Guesstimation", "Back of the envelope physics", the excellent "What If?" section of xkcd, and the lateral interview questions facing some job candidates: "how much of the world's water is contained in a cow?" are all examples.
 
Calculations in biology, such as the time it takes for a protein (foreground) to diffuse through an E. coli cell (background), are often subject to large uncertainties. Our approach and web tool allows us to track this uncertainty and obtain a probability distribution over possible answers (plotted).

We've built a free online calculator (Caladis -- calculate a distribution) that complements this approach by allowing one to take the uncertainty in one's estimates into account throughout a calculation. For example, what volume of CO2 is produced by our yearly driving? We could say that we cover 8000 miles per year "give or take" 1000 miles, and find that our car's CO2 emissions are between 100 and 150 grams per kilometre. Our calculator allows us to do the necessary conversions and sums while taking this possible variability into account -- doing maths with "probability distributions" describing our uncertainty. We no longer obtain a single (possibly inaccurate) answer, but a distribution telling us how likely any particular answer is -- in this case a rather concerning bell-shaped distribution between 1 and 2 tonnes which can be viewed here.

In the sciences, particularly in biology, measurements often have substantial uncertainties -- due to experimental error, natural variability in the system of interest, or both -- and so using distributions rather than single numbers in calculations allows us to understand and process more about the question of interest. "Back-of-the-envelope" calculations are certainly useful in biology but, owing to the uncertainties involved, one can trust one's estimates better if one has a smart envelope that takes that uncertainty into account.  We've written an accompanying paper (free here in Biophysical Journal) showing how to use our calculator -- in conjunction with the excellent Bionumbers online database, a collection of (often uncertain) experimental measurements in biology -- to make real biological calculations more powerful. Do have a go at using our calculator at www.caladis.org: it's user-friendly and there are lots of examples showing how it works! Iain and Nick [blog article also here]

ARTICLE: Fast inference about noisy biology

Efficient parametric inference for stochastic biological systems with measured variability


  • Understanding biology requires us to characterise the processes that go on inside our cells: we introduce a computational way to efficiently link models to observations, especially when these observations and processes are noisy.
Biology is a random and noisy world -- as we've written about several times before! This often means that when we try to measure something in biology -- for example, the number of a particular type of proteins in a cell, or the size of a cell -- we'll get rather different results in each cell we look at, because random differences between cells mean that the exact numbers are different in each case. How can we find a "true" picture? This is rather like working out if a coin is biased by looking at lots of coin-flip results.

Measuring these random differences between cells can actually tell us more about the underlying mechanisms for things like (to use the examples above) the cellular population of proteins, or cellular growth. However, it's not always straightforward to see how to use these measurements to fill out the details in models of these mechanisms. A model of a biological process (or any other process in the world) may have several "parameters" -- important numbers which determine how the model behaves (the bias of a coin, is an example, telling us what proportion of times we'll see heads). These parameters may include, for example, rates with which proteins are produced and degraded. The task of using measurements to determine the values of these parameters in a model is generally called "parametric inference". In a new paper, I describe a new and efficient way of performing this parametric inference given measurements of the mean and variance of biological quantities. This allows us to find a suitable model for a system describing both the average behaviour and typical departures from this average: the amount of randomness in the system. The algorithm I propose is an example of approximate Bayesian computation (ABC) which allows us to deal with rather "messy" data: I also describe a fast (analytic) approach that can be used when the data is less messy (Normally distributed).

The proposed efficient parametric inference process allows us to use data to make probabilistic statements about the details of models describing biology.

Parametric inference often consists of picking a trial set of parameters for a model and seeing if the model with those parameters does a good job of matching experimental data. If so, those parameters are recorded as a "good" set, otherwise, they're discarded as a "bad" set. The increase in efficiency in my proposed approach is due to the fact that we can perform a quick, preliminary check to see if a particular parameterisation is "bad", before spending more computer time on rigorously showing that it is "good". I show a couple of examples in which this preliminary checking (based on fast computation of mean results before using stochastic simulation to compute variances) speeds up the process by 20-50% on model biological problems -- hopefully allowing some scientists to grab a little more coffee time! This work is in the journal Statistical Applications in Genetics and Molecular Biology here, and you'll find the article (free) here. Iain [blog article also here]

ARTICLE: Exploring noise in cellular biology

The chaos within: Exploring noise in cellular biology


  • Biology requires exquisitely precise chemistry but exists in a very noisy world: we review how science is learning about how this randomness affects cell biology and how life has evolved to deal with it
We're used to thinking about machines as robust, hard-wearing objects made from solid materials like metal and plastics. If they crack, split or overheat they are liable to malfunction, and if we subject them to too much jostling and shaking we're asking for trouble. However, the biochemical machines responsible for keeping us alive work in a rather different world -- they're made from soft, organic materials, and contained in a disorganised bag (the cell) that is constantly shaken, bumping our machines against each other and other cellular inhabitants. How can the delicate processes required by living organisms take place in this chaotic environment? And how can scientific progress be made in such a tumultuous, unpredictable world?


Extrinsic factors can modulate the stability of essential, but noisy, cellular circuits: here we see that changes in transcription rate, provoked by differences in mitochondrial populations, affects the "landscape" underlying stem cell behaviour.

Iain recently wrote an article, targeted at a broad audience, looking at some of these questions. One of the most important cellular processes that has to take place in this chaotic world is that of 'gene expression': the interpretation of genetic blueprints which describe how to build cellular machinery, and the subsequent construction process. Gene expression can be likened to using a bad photocopier to copy books from a library that opens and closes randomly, then using these photocopies (which are prone to decay) to construct machines. This problematic environment gives rise to many medically important random effects, including bacterial resistance to antibiotics and differing responses to anti-cancer drugs. We are particularly interested in how fluctuating power supplies (see our other blog articles!) influence the cell's ability to produce these machines, and what effects this unreliable power has on medically important processes. The article -- available here and appearing in the expository magazine Significance here -- takes a look at how cellular noise arises, current techniques for its detection and analysis, and its influence on important biological phenomena. Iain [blog article also here]

ARTICLE: Taking the pulse of cellular power stations

Pulsing of Membrane Potential in Individual Mitochondria: A Stress-Induced Mechanism to Regulate Respiratory Bioenergetics in Arabidopsis


  • Mitochondria in plants sometimes switch off their membrane potential, which contributes to their ability to make energy for the cell: we characterise this "pulsing" and explore how it can be beneficial to plants under stress

We've just written an article in the journal Plant Cell about pulsing cellular power-stations and will motivate it by an analogy. Imagine we have a reservoir of water, and this water flows downhill through an outlet pipe, turning a turbine and producing energy. In this thought experiment, we're faced with a problem: the only way we can get water into our reservoir is by pumping it into the bottom of the reservoir. The higher up a reservoir is, the harder it is to pump water up there and the higher the risk of pumps overheating and getting damaged.

The problem can be solved by allowing the height of our reservoir to vary. If we lower our reservoir, it will become easier to fill, and the higher water pressure that arises from an increasingly filled reservoir will partly compensate for the fact that turbine-turning water will flow downhill from a lowered height, while allowing the pumps to relax and cool.

This model is a crude representation of mitochondria, the power stations of the cell, which use energy from respiration to create an energetic gradient across their membranes -- like a natural version of an AA battery. In our picture, this corresponds to the pumps feeding into our reservoir -- and in the cell, these pumps produce dangerous chemicals when they are overworked. The gradient they produce imbues protons with energy that is part electrical -- which we picture as the height of our reservoir -- and part chemical -- which we picture as the amount of water in our reservoir. These protons then flow through a protein complex -- the turbine -- to produce ATP, the universal cellular fuel.

An abstract representation (acrylic on canvas) of a single mitochondrion undergoing a 'pulse'. Its change in energy status is shown by the change in colour that we have also observed by microscopy using fluorescent sensors. Artist: Markus S

When mitochondria pump many protons, their "reservoirs" rise, with the increase in height forcing the pumps to work harder to pump water into the reservoir. This work produces dangerous chemicals which can damage the cell and the mitochondria themselves (called reactive oxygen species - they're what antioxidants try to combat). We have found a new mechanism by which this risk is decreased: if mitochondria are having to work hard, they "pulse", spontaneously lowering the height of their reservoir. This decreases the amount of work that the mitochondrial pumps have to do to fill the reservoir. The amount of turbine-turning energy per unit of water decreases, but as it becomes easier to fill the reservoir, more water gets pumped into it, partly compensating for the loss of height by an increase in volume. The pulsing process thus lowers the reservoir but fills it with more water, allowing the mitochondrial pumps to relax and reducing the production of dangerous chemicals.

We observed these pulses, spontaneous decreases of mitochondrial membrane potential, in Arabidopsis thaliana, a model plant species used in many biological contexts. Treating plant mitochondria with a variety of chemicals and observing the effects on pulsing, we deduced a biochemical mechanism by which pulsing occurs: a controlled influx of cations such as calcium ions into the mitochondrial matrix decreases membrane potential. We also found that pulsing is increased when plants face stressful environments: if they are suddenly heated, for example, or exposed to toxic chemicals. This novel mechanism may help explain some of the variability that our cellular engines exhibit and may be an important discovery in considering how mitochondria react to dangerous cellular conditions. You'll find the article here. Iain, Markus & Nick [blog article also here].