Showing posts with label parametric inference. Show all posts
Showing posts with label parametric inference. Show all posts

Friday, 23 April 2021

ARTICLE: How does tool use evolve in animals?

Data-Driven Inference Reveals Distinct and Conserved Dynamic Pathways of Tool Use Emergence across Animal Taxa, iScience 23 101245 (2020)

There are some wonderful examples of animals using tools. Octopuses block oyster shells open with coral; boxer crabs wave captive anemones for defense and food capture; captive dolphins use feather to wipe clean their aquarium windows. A while ago, we saw this excellent infographic in National Geographic, and got interested in these data. Do different families of animals learn to use tools in completely different ways? Or are there some general ("universal") principles behind how animals learn to use tools?

There is, of course, a big and fascinating literature on tool use, but we found rather few studies attempting a quantitative comparative analysis across bilaterian animals. To address other evolutionary questions, we've developed HyperTraPS (hypercubic transition path sampling), a statistical approach for learning the "pathways" of evolutionary processes. That is, which events occur before and after which other events in an evolving system? Does feature A always evolve before feature B? We used HyperTraPS to ask about the orderings with which different types of tool use appeared in animals. For example, do animals always learn to "poke" before they learn to "dig"? Do all animals learn tool use in the same order, first A then B then C..., or does it vary across species?

We found some answers that we think are quite interesting. There seem to be some quite deep similarities across animal species in how tool use evolves. Types of tool use like "affixing" and "throwing" are almost universally acquired early; types like "cutting" and "symbolising" are acquired late and rarely, only by primates. The environment and animal family influences the structure of these pathways: aquatic organisms seem to discover "waving" tool use relatively early, for example, and primates discover tools that "block" relatively late.

(A) The inferred pathways of tool use emergence across animals. The size of a blob gives the probability that that mode of tool use (on the horizontal axis) is acquired at that stage (on the vertical axis) of a species' discovery of tool use types. (B) Sample evolutionary pathways of tool use, with individual animal lineages illustrated at the positions corresponding to the modes of tool use they have discovered.

Of course, there's a lot of uncertainty about any analysis like this. Are we talking about wild or captured animals? What if we just haven't observed some types of tool use? We attempted to address several such questions with our analysis and showed that our overall results were quite robust with respect to these uncertainties. HyperTraPS fully describes the uncertainty in its outcomes, helping interpretability. We hope that our results help at least to suggest some possible principles and points for further investigation in this fascinating topic. You can read more in iScience here.

Monday, 15 July 2019

ARTICLE: Phenotypes and progression pathways in severe malaria

Precision identification of high-risk phenotypes and progression pathways in severe malaria without requiring longitudinal data
Iain G Johnston, Till Hoffmann, Sam F Greenbury, Ornella Cominetti, Muminatou Jallow, Dominic Kwiatkowski, Mauricio Barahona, Nick S Jones, Climent Casals-Pascual
npj Digital Medicine 2 63 (2019)





We recently published an article here in npj Digital Medicine using maths (including HyperTraPS) to learn more about severe malaria, a disease that kills over 400 000 people (mainly African children) a year. Severe malaria is challenging in the clinic because its symptoms and progress vary a lot from patient to patient. Our approach helps learn about this variability and identify high-risk patients and pathways. You can read our blog article about the paper on the npj Digital Medicine community blog here!



Thursday, 11 July 2019

ARTICLE: Getting to the root of the problem

Model selection and parameter estimation for root architecture models using likelihood-free inference
Clare Ziegler, Rosemary J. Dyson, Iain G. Johnston
J Roy Soc Interface (online, doi.org/10.1098/rsif.2019.0293 , 2019)

Roots bridge plants and soil, making vital contributions to crops, the environment, and fundamental biology. Because of this importance, understanding how roots grow under different conditions is a key scientific target. Experimental approaches to study roots can be challenging: being underground, it’s hard to observe root systems without perturbing them. Computer models can help here: we can simulate root growth and the “architecture” of root systems under lots of different conditions, without having to dig up and destroy real plants.


 
Observing roots growing underground is hard, but not impossible: here's a shot from our "minirhizotron" experiments using underground cameras to watch roots grow in an experimental woodland facility (see article here, and 3D version here!)

As computers have become more powerful, more and more sophisticated models for root growth and architecture have emerged. These simulation approaches typically take as input a set of parameters, and produce as output a model root system. These parameters are numbers describing, for example, the rates of root elongation, distances between lateral root branches, and so on – there may be dozens, or hundreds, of parameters in a sophisticated root model.

The output of a model depends strongly on these parameter values. So how can we choose the “right” ones? We may know some from experiments – for example, the widths of roots can be readily measured. But others may be less easy to observe. It is quite common to make educated guesses at these parameters, and see if the resulting root system “looks right”. This approach has a few issues – it can be subjective, and doesn’t give us information on how flexible our guesses are. For example, is a growth rate of 0.1cm per day just as likely as 0.5cm per day, or 0.02cm per day? And how can we tell if one version of a model does “better” than another, and is more supported by real observations?

In a new paper here in Journal of the Royal Society Interface, we propose a platform to provide answers to these questions, using so-called “approximate Bayesian computation” or ABC. This is a way of learning which parameter values and models are most compatible with observed data, by running many simulations with many different choices, and comparing the output of each choice to our observations using specific criteria. This replaces the subjective “looks right” and explores a wide set of parameterisations, allowing us to learn what ranges of values are most likely given our data. We can also use ABC to compare different mechanisms for root growth, finding which is most supported by observation. This helps us gain scientific insight and ensures that the outputs of our models can be more reliably intepreted.


Overview of our approach. Using ABC allows us to identify governing parameters and mechanisms for root growth that are most supported by real observations.

We tested our ABC approach with synthetic observations from models of thale cress and narrowleaf lupin, confirming that we can recover the parameter values we put in. We then used real thale cress plants (wild and mutant) to show that our platform distinguishes genetically different plants and identifies most-likely parameters and model structures for real root growth. We used the platform to select models for growth and branching, showing how it can be used to compare existing models from the literature. We hope that this approach can be used to further help improve the interpretability and rigour of plant modelling and simulation! Iain and Clare

Saturday, 22 September 2018

ARTICLE: Time marches on -- mitochondria, ageing, and disease

Burgstaller, J.P., Kolbe, T., Havlicek, V., Hembach, S., Poulton, J., Piálek, J., Steinborn, R., Rülicke, T., Brem, G., Jones, N.S. and Johnston, I.G. Large-scale genetic analysis reveals mammalian mtDNA heteroplasmy dynamics and variance increase through lifetimes and generations. Nature communications2488 (2018)

DNA in mitochondria, the powerhouses of the cell, is passed down from mother to child. But there are many mitochondria in each cell, and these mitochondria may have different genetic features. If a mother carries a mixture of mitochondrial DNA (mtDNA) types, this can make it hard to say which features their children will inherit. For mothers carrying a disease-causing mtDNA mutation, this makes family planning and clinical therapies challenging.


In particular, the role of a mother's age has long been a mystery. Is the probability of a child inheriting a particular mtDNA feature higher when mothers are younger or older? An answer to this question could help plan clinical strategies to improve fertility and prevent the inheritance of deadly mitochondrial disease.


To address this, we worked with our excellent collaborators with a combination of maths, statistics, and experiment. Our collaborators used cutting-edge technology to reveal the mixtures of mtDNA in the egg cells of mother mice at a wide range of ages, and in the litters of offspring the mothers produced. This experimental work was the largest-scale study of mammalian mtDNA that we're aware of, involving thousands of observations throughout lifetimes and between generations. In concert, we developed a mathematical model describing the changes to, and inheritance of, mtDNA from mother to offspring. We combined the model and data to learn how different biological processes affect mtDNA through and between generations.



Cells contain populations of mitochondria, and these populations change over time. In European mice, we observed how variability in these populations evolves as mammals age and reproduce. We found that older mother have more varied mitochondria and pass this variance on to their offspring -- of central importance in the inheritance of genetic disease. 

We found that the variability of mtDNA dramatically increased as mothers aged. This means that the probability of inheriting more extreme -- both lower and higher -- levels of a genetic feature increases for older mothers. We also found that different mtDNA mixtures were inherited in different ways - with some mtDNA types favoured for inheritance and some disfavoured. We used our findings to create a way to predict how the risk that offspring would inherit disease-causing mtDNA features changes over time. Moving forward, we're aiming to harness these powerful ways of using large datasets to describe and predict the dynamics of mtDNA inheritance in humans, and to learn what it is about these mtDNA types that predicts their evolution across generations. You can read the article for free in Nature Communications here.

ARTICLE: How cells adapt to progressive increase in mitochondrial mutation

Aryaman, J., Johnston, I.G. and Jones, N.S. Mitochondrial DNA density homeostasis accounts for a threshold effect in a cybrid model of a human mitochondrial disease. Biochemical Journal474 4019 (2017).

Mitochondria produce the cell's major energy currency: ATP. If mitochondria become dysfunctional, this can be associated with a variety of devastating diseases, from Parkinson's disease to cancer. Technological advances have allowed us to generate huge volumes of data about these diseases. However, it can be a challenge to turn these large, complicated, datasets into basic understanding of how these diseases work, so that we can come up with rational treatments.


We were interested in a dataset (see here) which measured what happened to cells as their mitochondria became progressively more dysfunctional. A typical cell has roughly 1000 copies of mitochondrial DNA (mtDNA), which contains information on how to build some of the most important parts of the machinery responsible for making ATP in your cells. When mitochondrial DNA becomes mutated, these instructions accumulate errors, preventing the cell's energy machinery from working properly. Since your cells each contain about 1000 copies of mitochondrial DNA, it is interesting to think about what happens to a cell as the fraction of mutated mitochondrial DNA (called 'heteroplasmy') gradually increases.  We used maths to try and explain how a cell attempts to cope with increasing levels of heteroplasmy, resulting in a wealth of hypotheses which we hope to explore experimentally in the future.





The central idea arising from our analysis of this large dataset is that cells seem to attempt to maintain the number of normal mtDNAs per cell volume as heteroplasmy initially increases from 0% mutant. We suggest they do this by shrinking their size. By getting smaller, cells are able to reduce their energy demands as the fraction of mutant mtDNA increases, allowing them to balance their energy budget and maintain energy supply = demand. However, cells can only get so small and eventually the cell must change its strategy. At a critical fraction of mutated mtDNA (h* in the cartoon above), we suggest that cells switch on an alternative energy production mode called glycolysis. This causes energy supply to increase, and as a result, cells grow larger in size again. These ideas, as well as experimental proposals to test them, are freely available in the Biochemical Journal "Mitochondrial DNA Density Homeostasis Accounts for a Threshold Effect in a Cybrid Model of a Human Mitochondrial Disease". Juvid, Iain and Nick

ARTICLE: How plants decide when to germinate

Topham, A.T., Taylor, R.E., Yan, D., Nambara, E., Johnston, I.G. and Bassel, G.W. Temperature variability is integrated by a spatially embedded decision-making center to break dormancy in Arabidopsis seeds. PNAS 114 6629 (2017)

A plant's choice to germinate is one of the most important decisions in the world. If it is made too soon, the plant may be damaged by harsh winter conditions; if too late, the plant may be outcompeted, and crop yields may be lower. If crops in a field make the decision at different times, there is more room for weeds to grow and pests to take over. 


In a recent study, we combined mathematical modelling with several neat experiments to identify sets of cells that make this germination choice in a much-studied plant called thale cress (Arabidopsis thaliana), and have learned how it makes decisions based on the plant's environment.



Two views of the plant embryo from laser microscopy, highlighting cells where different components of the germination control machinery are expressed. The background shows the "attractor basins" in a mathematical description of the germination decision: horizontal and vertical axes give the levels of two hormones ABA and GA, the blue region corresponds to dormant seeds and the red region to germination. 

This germination circuitry functions through a circuit of chemical stimuli and responses. Using laser microscopy, we found that different parts of this circuit exist in different parts of the plant embryo -- and that the separation of these parts is central to how the brain functions. We used mathematical modelling to show that communication between separated elements of the germination circuitry controls the plant's sensitivity to its environment. Following this theory, we used a mutant plant where cells were more chemically linked -- essentially enhancing communication between circuit elements -- to show that germination depends on these intra-cellular signals.


The separation of circuit elements allows a wider palette of responses to stimuli. It's like the difference between reading one critic's review of a film four times over, or amalgamating four different critics' views before deciding to go to the cinema. Our mathematical theory predicted that more plants would germinate when exposed to varying environments -- like three short pulses of cold -- than constant environments -- like one long cold period. We tested this theory in the lab and found exactly this behaviour.


Next, the hope is to learn about the germination brain in other plants and crops, and to show how our new knowledge of the germination machinery can be used to enhance and synchronise germination in crops. You can read the paper for free in the journal PNAS here. Iain

Sunday, 21 May 2017

ARTICLE: A healthy dose of mathematics

Toward Precision Healthcare: Context and Mathematical Challenges
C Colijn, N Jones, IG Johnston, S Yaliraki, M Barahona
Frontiers in Physiology 8 136 (2017)
  • The continuing explosion of available biomedical data will help us tailor and optimise therapies for individual patients; we are designing new maths and statistics to help this process and to include social and other data into an overarching "precision healthcare" approach.
Our research combines tools from maths and statistics with biological data to learn more about the biological world. An exciting, growing, and much-discussed branch of science -- precision medicine -- is a specific instance of this idea. The vision of precision medicine is to use the expanding volume of data that's emerging from medicine and biology to tailor and optimise medical therapies for individual patients, making the therapies as effective as possible. This idea isn't new -- we are well aware, for example, that an individual's blood type dictates which blood transfusions they can successfully receive. But precision medicine is a much bigger picture, potentially taking into account large amounts of genetic, environmental, dietary, and other features to identify the optimal treatment for a disease -- for example, tailoring chemotherapy treatments to match the genetic specifics of a particular cancer case.

Dealing with these large and diverse datasets will need new mathematical and statistical approaches, built with an ongoing link to clinical practice. At the same time, we're interested in expanding the idea of precision medicine to include the "big data" that's increasingly available about individuals' social and logistic contexts. Social networks can dictate how diseases spread -- and how knowledge and views about therapies, vaccines, and other medically pertinent ideas are transmitted and shaped from person to person. A person's home region determines the genetic structure of local people who may act as donors. We're looking at the idea of "precision healthcare" -- using new maths and statistics to optimise healthcare strategy, not just individual therapies, in the light of large-scale datasets.





One aspect of precision healthcare we'll be exploring is exploring how progressive diseases -- those that involve the accumulation of symptoms over time -- develop in patients, using transition networks like those above to model "disease spaces" and find pathways in those spaces.

We're excited to be part of a new initiative -- the Centre for the Mathematics of Precision Healthcare -- involving six parallel and related research projects that align with this goal. Some of our previous work -- for example, estimating social structures of big UK cities to explore the challenges that genetic diversity poses to gene therapies for mtDNA disease -- already has a precision healthcare feel. In a new review paper (available for free) we discuss this and other examples of past and future work that we hope will contribute to the precision healthcare goal, along with some key ideas and context for the initiative. Iain

Friday, 29 January 2016

ARTICLE: Go green -- recycle mitochondria

A novel quantitative assay of mitophagy: Combining high content fluorescence microscopy and mitochondrial DNA load to quantify mitophagy and identify novel pharmacological tools against pathogenic heteroplasmic mtDNA

  • Mitophagy degrades mitochondria, and likely plays important roles in the cell's responses to mitochondrial disease, but is hard to measure and thus poorly understood: we propose new ways of measuring mitophagy and use them to explore drugs that may help change damaged mitochondrial populations
Mitochondria, as we've written about before, are important entities in our cells that produce energy and take part in many other vital processes. Mitochondrial DNA (mtDNA), inherited from our mothers, contains instructions on how to build important mitochondrial machinery. MtDNA is sometimes mutated, leading to problems with our mitochondria. How do our cells cope?

Mitophagy (from mito-(chondria) and -phagy (eating)) is a process by which cells degrade and recycle mitochondria, allowing dysfunctional mitochondria to be removed and replaced. Mitophagy is one of a number of cellular mechanisms that maintain a healthy population of mitochondria, and appears to play a central role in determining the inheritance and evolution of mtDNA over our lifetimes. However, our understanding of mitophagy is limited because it is hard to observe.

In a recent and epically-titled paper in Pharmacological Research here, we explore two different approaches for measuring mitophagy in cells. The first is physical. We used chemicals to make mitochondria glow red, and autophagosomes (the cellular machines responsible for the degradation of mitochondria) glow green. We then used a microscope to examine large numbers of cells and recorded how often red (mitochondria) and green (autophagosomes) were seen together, which we took to imply that mitophagy may be occurring. We confirmed that various drugs and chemicals known to affect mitophagy had the expected effects on this estimate of mitophagy, and that perturbing ATG7 (an essential part of the autophagic machinery) sustantially reduced our observed mitophagy levels.

We also subjected cells to stress by growing them with a less plentiful supply of energy. We found that this energy stress increased the amount of mitophagy (perhaps as cells struggle to make the very best of their mitochondrial populations). We also found that mitophagy broadly decreased in cells from older people, and was increased in cells from people carrying an mtDNA disease (negatively affecting mitochondrial functionality).

The second approach is genetic. In cells from patients with mtDNA disease, some mtDNA is normal and some is mutated -- we used genetic tools to measure the proportion of mutant mtDNA in cells. We observed that when we stressed patients' cells, levels of mutant mtDNA decreased while our physically observed measure of mitophagy increased, supporting a picture in which mitophagy removes dysfunctional mitochondria when energy output is of central importance. We also found evidence for undirected mitophagy, where mtDNA copy number is depleted with no preference for mutant or wildtype.

Observing the colocalisation of autophagosomes (green) and mitochondria (red), as well as the proportion of mutant mtDNA (white stars), allows a bilateral characterisation of mitophagy. The patterns of changes in these observations tell us about how drug treatments and different environments change mitochondrial populations.

The physical and genetic approaches give us two largely independent means to estimate mitophagy, placing our understanding of this vital process on a solid analytical foundation. We used these tools to assess the effects of various drugs on mitophagy, allowing us to characterise the effects of drugs like metformin (inhibiting mitophagy) and phenanthroline (inducing undirected mitophagy) in unprecedented detail and facilitating more precise statements about their utility in clinical contexts. Iain


Wednesday, 27 January 2016

ARTICLE: Warburg Ensemble

Monitoring Intracellular Oxygen Concentration: Implications for Hypoxia Studies and Real-time Oxygen Monitoring

  • Cancer cells vary in how they produce their energy: we make progress understanding this variability, which may eventually help scientists design better therapies.
Cells can produce energy through several processes. We'll consider two – process "O" (for "oxidative phosphorylation"), and process "G" (for "glycolysis"). "O" uses oxygen, and harnesses the cell's mitochondria to produce energy. "G" does not use oxygen and produces energy without directly using mitochondria.

Healthy cells use both “O” and “G”, but cancer cells are often observed to rely on "G" much more. The shift away from "O+G" towards just "G" in cancer is often called the "Warburg effect", after Otto Warburg, who wrote about the shift in the 1950s. It remains unclear, however, whether the Warburg effect applies to all cancer cells under all conditions, or if different cells and different environments experience different shifts. This is important because understanding how cancer cells get their energy -- and, more generally, what changes occur in cancer cells compared to healthy cells -- may allow us to design therapies that challenge cancer cells while leaving healthy cells undamaged.

We used some fancy modern technology (focussed around the MitoXpress-Intra probe) to measure the difference between oxygen levels within a cell and oxygen levels in the cell's environment. We developed a mathematical way of producing "calibration curves", directly linking the observed MitoXpress behaviour to oxygen concentrations. If cells are using "G" alone, these levels are similar, as no oxygen is being consumed by the cells. If cells are also using "O", oxygen levels within cells should be rather lower than in their environment.

We found that two different cancer cell lines (with the rather jargon-y names "RD" and "U87MG") behaved surprisingly differently. When grown on glucose, U87MG looks quite "G", with oxygen levels within cells similar to those in the environment (e.g. 17.1% in cells, 18% outside). RD looks much more "O+G", with substantial differences between in-cell and outside-cell oxygen levels (e.g 13.2% in cells, 18% outside). Importantly, these findings were reproduced across a range of environmental oxygen levels (18% to 5%), modelling the range of conditions that cancer cells experience in tumours in the body. The two cancer cell lines thus seem to produce their energy in rather different ways, underlining that the Warburg effect is not an invariant across all cancers, and that treatments may be improved by taking this into account. We also showed that treating a different cancer cell line ("786-0") with phenformin, a drug inhibiting mitochondria, shifts cells away from "O+G" to "G", and that this shift can be monitored in real time with MitoXpress.

Different cancer cell lines (U87MG and RD) produce energy through different pathways, engaging more “G” (glycolysis) or “O” (oxidative phosphorylation). “O” uses oxygen (O2), lowering oxygen levels in cells compared to their environment. The different balance of “G” and “O” in different cases is important for understanding the heterogeneity of cancer.

Our paper appears in a book with the catchy title "Oxygen Transport to Tissue XXXVII", associated with the journal Advances in Experimental Medicine and Biology. You can get a sneak peek here and we'll update with a link when possible. Iain

ARTICLE: How evolution deals with mitochondrial mutants (and how we can take advantage)

Stochastic modelling, Bayesian inference, and new in vivo measurements elucidate the debated mtDNA bottleneck mechanism

  • Disease-causing mutant mtDNA is inherited through a complicated process: we use maths and statistics to shed light on this process and suggest possible therapeutic strategies to address disease inheritance and onset
Our mitochondrial DNA (mtDNA) provides instructions for building vital machinery in our cells. MtDNA is inherited from our mothers, but the process of inheritance -- which is important in predicting and dealing with genetic disease -- is poorly understood. This is because mitochondrial behaviour during development (the process through which a fertilised egg becomes an independent organism) is rather complex. If a mother's egg cell begins with a mixed population of mtDNA -- say with some type A and some type B -- we usually observe hard-to-predict mtDNA differences between cells in the daughter. So if the mother's egg cell starts off with 20% type A, egg cells in the daughter could range (for example) from 10%-30% of type A, with each different cell having a different proportion of A. This increase in variability, referred to as the mtDNA bottleneck, is important for the inheritance of disease. It allows cells with higher proportions of mutant mtDNA to be removed; but also means that some cells in the next generation may contain a dangerous amount of mutant mtDNA. Crucially, how this increase in variability comes about during development is debated. Does variability increase because of random partitioning of mtDNAs at cell divisions? Is it due to the decreased number of mtDNAs per cell, increasing the magnitude of genetic drift? Or does something occur during later development to induce the variability? Without knowing this in detail, it is hard to propose therapies or make predictions addressing the inheritance of disease.

We set out to answer this question with maths! Several studies have provided data on this process by measuring the statistics of mixed mtDNA populations during development in mice. The different studies provided different interpretations of these results, proposing several different mechanisms for the bottleneck. We built a mathematical framework that was capable of modelling all the different mechanisms that had been proposed. We then used a statistical approach called approximate Bayesian computation to see which mechanism was most supported by the existing data. We identified a model where a combination of copy number reduction and random mtDNA duplications and deletions is responsible for the bottleneck. Exactly how much variability is due to each of these effects is flexible -- going some way towards explaining the existing debate in the literature.  We were also able to solve the equations describing the most likely model analytically. These solutions allow us to explore the behaviour of the bottleneck in detail, and we use this ability to propose several therapeutic approaches to increase the "power" of the bottleneck, and to increase the accuracy of sampling in IVF approaches.

A "bottleneck" acts to increase mtDNA variability between generations. But how is this bottleneck manifest? Our approach suggests that a combination of copy number reduction (pictured as a "true" copy number bottleneck), and later random turnover of mtDNA (pictured as replication and degradation), is responsible.

Our excellent experimental collaborators, lead by Joerg Burgstaller, then tested our theory by taking mtDNA measurements from a model mouse that differed from those used previously and which, could in principle have shown different behaviour. The behaviour they observed agreed very well with the predictions of our theory, providing encouraging validation that we have identified a likely mechanism for the bottleneck. New measurements also showed, interestingly, that the behaviour of the bottleneck looks similar in genetically diverse systems, providing evidence for its generality. You can read about this in the free (open-access) journal eLife here. Iain and Nick [blog article also here]

ARTICLE: Great technological power, great statistical responsibility

Multiple hypothesis correction is vital and undermines reported mtDNA links to diseases including AIDS, cancer, and Huntingdon’s

  • Several papers perform incorrect and misleading statistical analyses in seeking links between mtDNA and cancer: these statistical issues must be corrected before scientific and policy progress can be made from these investigations
Biologists often report a result as a "significant" sign of exciting new science if there is less than a 1-in-20 chance that the result they observe could have emerged by chance from boring old science. This is silly (although we do it too!) -- by contrast, for example, physicists require less than a 1-in-3,500,000 chance. But this post won't discuss too many problems with this state of affairs -- that is done admirably elsewhere.

The problem can be compounded when scientists take lots of measurements. Say we take 50 measurements of a boring old system, and every time we see something that has less than a 1-in-20 chance of appearing in a boring old system, we call it "significant". We're playing the odds 50 times, so we expect to see 1-in-20 results appear around 2 or 3 times; just as if we roll a dice 50 times, we'd expect to roll a good few sixes. If we call every 1-in-20 result "significant" without accounting for the fact that we've looked at lots of measurements (and are thus more likely to see 1-in-20s by chance), we are in danger of reporting exciting new science when in fact the boring old science has been true all along.

There are lots of ways of doing this accounting, but a series of papers that have been recently published linking mtDNA to diseases have made no attempt to do it. Generally, these papers look at the mtDNA of people without the disease and the mtDNA of people with the disease. If any mtDNA features appear more in the people with the disease, the paper calculates the chance of that difference occurring in the boring old picture (in which there is no link between the mtDNA feature and the disease). If they drop below the 1-in-20 mark, they report an exciting new link between that feature and the disease. But they test dozens of features and never account for this multiple testing -- so, as above, we'd expect them to see "significant" results emerging just by chance. In a paper in Mitochondrial DNA here (free here) I show, by creating artificial data, that this problem is rife, that most of these reported links are spurious, and that scientists really need to be more responsible, before their flawed analysis starts to misguide health policy and medicine.

The top graph shows how the probability of seeing a 1-in-20 occurrence (p < 0.05 in the jargon), when in fact there is nothing new and exciting to report, increases as a scientist investigates more things. If an experiment consists of one test, then a 1-in-20 occurrence indeed has a 1-in-20 probability (0.05). But as soon as we do more tests, the chance of seeing at least one 1-in-20 occurrence starts to increase, as we are "playing the game" more times. If we do 6 tests there is a 0.27 probability -- between a 1-in-4 and 1-in-3 chance -- that we will see at least one 1-in-20 event. This is illustrated below, where we have six dice and think some of them may be unfair. We roll each one five times and count the number of 6s. One of them comes up 6 three times -- the chance of this happening for one fair die is less than 1-in-20. But because we've looked at six dice, we should be less surprised to see this rare event, because we've looked at more events in total. We need more evidence to claim that this die is unfair.

This quick note only represents the tip of the iceberg. MtDNA studies are often statistically unsound; statistical misdemeanours in biomedical studies are so common that most published research is wrong; scientists increasingly focus on the 1-in-20 chance as opposed to the size and importance of the effect they're measuring; the majority of hallmark papers in vital fields like cancer science are unreproducible (though this last point may have other causes than statistical problems). The 1-in-20 idea was only ever meant to be a step in identifying interesting scientific avenues, not the final measure of scientific truth. This is a big, and growing, problem! Iain

(For accessibility I have used "exciting", "boring", and "1-in-20" instead of their usual, more technical labels; they of course are usually called the "alternative hypothesis", "null hypothesis", and "p < 0.05" respectively).

ARTICLE: Evolutionary competition within our cells: the maths of mitochondrial DNA

mtDNA Segregation in Heteroplasmic Tissues Is Common In Vivo and Modulated by Haplotype Differences and Developmental Stage


  • MtDNA mixtures in cells arise through mutation and gene therapies: we show that different types of mtDNA usually proliferate at different rates, which suggests ways that therapies could be made more efficient.
Women may carry mutated copies of mitochondrial DNA (mtDNA) -- a molecule that describes how to build important cellular machinery relating to cellular energy supply. If this mutant mtDNA is passed on to that woman's child, the child may develop a mitochondrial disease, which are often degenerative, fatal, and incurable.

Joerg created mice that contained two types of mtDNA -- here illustrated as blue (lab mouse mtDNA) and yellow (mtDNA from a mouse from a wild population). We used several different wild mice from across Europe to represent the mtDNA diversity one may find in a human population. We found that throughout a mouse's lifetime, one mtDNA type often outcompetes another (here, yellow beats blue), with different patterns across different tissues.

Amazing new therapies potentially allow a carrier mother A and a father B to use another woman C's egg cells to conceive a baby without much of mother A's mtDNA being present. The approach involves taking nuclear DNA content from A and B (so that most of the child's features are inherited from the true mother and father), and placing it into C's egg cells, which contain a background of healthy mtDNA. You can read about, what are misleadingly called, three-parent babies here.

Something that is less discussed is that, in this process, a small amount of A's mutant mtDNA can be "carried over" into C's cell. If this small amount remains small through the child's life, there is no danger of disease, as the larger amount of healthy C mtDNA will allow the child's cell to function normally. We can think of the resulting situation as a competition between A and C -- if A and C are evenly matched, the small amount of A will remain small; if C beats A, the small amount of A will disappear with time; and if A beats C, the small amount of A will increase and may eventually come to dominate over C.

Until recently it has been fair to assume that A and C are always about evenly matched (unless something is drastically different between A or C). However, evidence for this idea was based on model organisms in laboratories, which do not have the same amount of genetic diversity as found in human populations. Our collaborator Joerg addressed this by capturing wild mice from across central Europe, selecting a set that showed a comparable degree of genetic diversity to that expected in a human population. He used these, with our modelling and mathematical analysis, to show that pronounced differences between A and C often exist, and are more likely in more diverse populations. The possibility that A beats C, and mutant mtDNA comes to dominate the child's cells, therefore cannot be immediately discounted in a diverse population. We propose "haplotype matching" -- ensuring that A and C are as similar as possible -- to ameliorate this potential risk. It's open as to whether one can generalize from observations in mice to people and it's also open as to whether our conclusions, which used lab-mice as parent A (which are not entirely typical creatures) of necessity generalize to other non-lab mouse types.

Our mathematical approach also allowed us to explore, in detail, the dynamics by which this competition within cells occurs. We were able to use our data rather effectively by having a statistical model that allowed us to reason jointly about a range of data sets. We found that the degree to which one population of mtDNA beat the other depended on how genetically different they were.  We found that different tissues were like different environments: some favouring C over A and some vice-versa. This is perhaps surprising to some as this evolution in the proportions of different genetic species is not something we imagine occurring inside us, during our lives, and as something that might differ between our organs. We found several different regimes, where the strength of competition changes with time and as the organism develops: when our cells are multiplying faster they show a more marked preference for one of the species. We've shown our results to the UK HFEA in its ongoing assessment of these therapies, and you can read, for free, about our work in the journal Cell Reports here. Iain, Joerg, Nick [blog article also here].

ARTICLE: Fast inference about noisy biology

Efficient parametric inference for stochastic biological systems with measured variability


  • Understanding biology requires us to characterise the processes that go on inside our cells: we introduce a computational way to efficiently link models to observations, especially when these observations and processes are noisy.
Biology is a random and noisy world -- as we've written about several times before! This often means that when we try to measure something in biology -- for example, the number of a particular type of proteins in a cell, or the size of a cell -- we'll get rather different results in each cell we look at, because random differences between cells mean that the exact numbers are different in each case. How can we find a "true" picture? This is rather like working out if a coin is biased by looking at lots of coin-flip results.

Measuring these random differences between cells can actually tell us more about the underlying mechanisms for things like (to use the examples above) the cellular population of proteins, or cellular growth. However, it's not always straightforward to see how to use these measurements to fill out the details in models of these mechanisms. A model of a biological process (or any other process in the world) may have several "parameters" -- important numbers which determine how the model behaves (the bias of a coin, is an example, telling us what proportion of times we'll see heads). These parameters may include, for example, rates with which proteins are produced and degraded. The task of using measurements to determine the values of these parameters in a model is generally called "parametric inference". In a new paper, I describe a new and efficient way of performing this parametric inference given measurements of the mean and variance of biological quantities. This allows us to find a suitable model for a system describing both the average behaviour and typical departures from this average: the amount of randomness in the system. The algorithm I propose is an example of approximate Bayesian computation (ABC) which allows us to deal with rather "messy" data: I also describe a fast (analytic) approach that can be used when the data is less messy (Normally distributed).

The proposed efficient parametric inference process allows us to use data to make probabilistic statements about the details of models describing biology.

Parametric inference often consists of picking a trial set of parameters for a model and seeing if the model with those parameters does a good job of matching experimental data. If so, those parameters are recorded as a "good" set, otherwise, they're discarded as a "bad" set. The increase in efficiency in my proposed approach is due to the fact that we can perform a quick, preliminary check to see if a particular parameterisation is "bad", before spending more computer time on rigorously showing that it is "good". I show a couple of examples in which this preliminary checking (based on fast computation of mean results before using stochastic simulation to compute variances) speeds up the process by 20-50% on model biological problems -- hopefully allowing some scientists to grab a little more coffee time! This work is in the journal Statistical Applications in Genetics and Molecular Biology here, and you'll find the article (free) here. Iain [blog article also here]

ARTICLE: Inferring the evolutionary history of photosynthesis : C 4 yourself


Phenotypic landscape inference reveals multiple evolutionary paths to C4 photosynthesis

  • Some plants have evolved efficient photosynthesis, but important crops including rice have not: we use maths and statistics in conjunction with biological data to understand this evolution, with a view to repeating it artificially in crops to increase food production

Biological evolution is a complex, stochastic process which dictates fundamental properties of life. Our understanding of evolutionary history is severely limited by the sparsity of the fossil record: we only have a handful of fossilised snapshots to infer how evolution may have progressed throughout the history of life. Many physicists and mathematicians have attempted theoretical treatments of the process of evolution, using varying degrees of abstraction, in order to provide a more solid quantitative foundation with which to study this complex and important phenomenon, but the predictive power of these theoretical models, and their ability to answer specific biological questions, is often questioned.

We recently focussed on one remarkable product of evolution in plants: so-called "C4 photosynthesis". C4 consists of a complex set of changes to the genetic and physiological features which have evolved in some plants and act to increase the efficiency of photosynthesis. This complex set of changes has evolved over 60 times convergently: that is, plants from many different lineages independently "discover" C4 photosynthesis through evolution. We were interested in the evolutionary history of how these discoveries occurred -- both motivated by fundamental biology and the possibility of "learning from evolution" and using information about the evolution of C4 to design more efficient crop plants.


C3 and C4 plants differ in several physical and genetic ways (leaves and cells either side). We picture evolution as progressing along paths over a hypercube connecting these states (grey lines) -- some paths will give rise to intermediate species matching those we really observe (red and blue points). We can calculate how likely each path is and thus reconstruct evolutionary history.
 
To this end, we modelled the evolution of C4 as a pathway through a space containing many different possible plant features. The pathway starts at C3 -- the precursor to C4 -- and progressively takes steps in different directions, acquiring one-by-one the features that sum up to C4 photosynthesis. Using a survey of plant properties from across the wide scientific literature, we identified which intermediate states these pathways were likely to pass through, given observed properties of plants that currently possess some, but not all, C4 features. We were then able to use a new inference technique to predict the ordering in which these likely pathways traverse the evolutionary space. We showed that this approach worked by both successfully inferring the known evolutionary steps in synthetic datasets and correctly predicting previously unknown properties of several plants, which we verified experimentally. Our (open access) paper is here and there's a less technical summary and commentary here. Our approach showed that C4 photosynthesis can evolve through a range of distinct evolutionary pathways, providing a potential explanation for its striking convergence. Several of these different pathways were made explicitly visible when we examined the inferred evolutionary histories of different plant lineages -- different families are likely to have converged on C4 through different evolutionary routes. Furthermore, the most likely initial steps towards C4 photosynthesis are surprisingly not directly related to photosynthesis, being solutions to different biological challenges, but also providing evolutionary "foundations" upon which the machinery of C4 can evolve further. We hope that the recipes for C4 photosynthesis that we have inferred find use in efficient crop design, and anticipate our inference procedure being of use in the study of other specific biological questions regarding evolutionary histories. Iain [blog article also here]