Monday, 29 February 2016

Optimal odour receptors



The sense of smell is the ability to detect molecules of different chemicals in the air. This happens via an array of receptor molecules in your nose that bind to odour molecules and send signals to your brain. A simple design would have one receptor for each molecule. However it’s more complicated than that. Humans can detect over 2100 different molecules and can tell the difference between mixtures of up to 30 different molecules, using a much smaller number of distinct receptors. Receptors therefore need to respond to more than one different molecule each and the brain must put together all of the signals form the receptors to identify the odour. A recent paper by David Zwicker, Arvind Murugan and Michael P. Brenner discusses the optimal setup of receptors from an information-theoretic perspective.

One important point to note is that different odours are present in the natural environment with different frequencies and this ought to be taken into account when designing your receptor array. It is no use having an array that very accurately distinguishes between two very rare odours if it can’t distinguish between two very common odours. This is made more precise by Laughlin’s principle which says that an optimal sensor (one that conveys the most information) should be such that all possible outputs are equally likely in the natural environment.

Zwicker, Murugan and Brenner construct a simple model of the response of the receptor array to different mixtures of odour molecules in which the receptors respond proportionally to the concentrations of the molecules. Each receptor has a different sensitivity to each of the molecules and it transmits a binary signal that is 1 if the excitation is above a threshold and 0 if it is below. They find that there are two principles that are relevant to maximising the information:

(1) Any given receptor should be active half the time, which maximises the amount of information provided by that receptor in isolation.

(2) Each pair of receptors should have as uncorrelated (non-matching) a response as possible, reducing the redundancy between them.

They then go on to discuss to what extent these principles can be satisfied in general and which parameters of their model give the maximum information. The most interesting outcome is a fundamental trade-off between being effective at identifying the presence or not of as many molecules as possible, and being able to estimate the concentrations of molecules accurately. This trade-off arises because the best way to identify as many molecules as possible is for each to activate only a single receptor, whereas to estimate concentration each molecule activate a number of receptors with different thresholds.

This paper is a good example of one of the ideas in in William Bialek’s paper that Tom McGrath wrote about in a previous post. The authors are assuming that this biological process is operating near to the optimum that physical limits allow and then they  are trying to find the parameters of the system that correspond to these limits.  One criticism of the paper might be that it takes no account of the fact that some odours a more important than others, even if they are very rare. For example, accurately identifying the smell of a poisonous substance that an animal encounters occasionally might be more important than distinguishing between the smells of two good food sources that are encountered more frequently.

Wednesday, 20 January 2016

New contributors

This term I'll be encouraging some of the students I work with here in Imperial to contribute to this blog; Tom McGrath has corageously stepped up to take a first swipe with a discussion of a recent review from Bill Bialek on physics in biology.

Thoughts on William Bialek's "Perspectives on theory

In modelling almost anything in biology the complexity of the process leads almost inevitably to an explosion of free parameters whose tuning may give very different results.While this is often explained away as inevitable due to the rich diversity of the natural world, for a physicist there is something unsettling about a profusion of parameters, coupled with a hope that they could be obtained instead by resorting to some underlying principle. The question is: what principle (or principles) might be at work here? A recent post on the arXiV by Professor William Bialek at Princeton and CUNY summarises a discussion on this topic at the Simons Foundation 2014 Theory in Biology workshop.

The main goal of this paper is to explore possible guiding principles that allow us to recover the correct spot in parameter space, or render precise tuning of parameters unnecessary. A second topic is a pointed but valid critique of how theory (whether implicit or in an explicit mathematical form) is currently considered in the biological community, which I won't discuss here - if you're interested look in sections II - IV and IX of the paper. The three classes of principles explored in the paper are: functional behaviours emerging as 'robust' properties of biological systems without precise tuning (section VI), optimality arguments based on evolution (section VII), and emergent phenomena from large interacting systems (section VIII). In the interests of remaining concise I'll miss out a lot of the interesting examples from the paper and concentrate instead on one or two of the main examples in each section - the paper itself is well worth a look for these alone.

In Section VI the main emphasis is on robust behaviours - properties where entire regions of parameter space will produce similar results. In this case the main example is spike train production in neurons. Varying the copy number of protein channels in neurons produces neurons with qualitatively different spike train characteristics including silence, single, double, and triple spike bursts and rapid repeated spiking. This resolves a continuous parameter space into a set of distinct objects which can be combined to form a functional system; by adjusting the copy numbers of any particular neuron in a network it can be robustly changed from one type to another. This suggests a view of neurons as more like building blocks than objects for individual investigation (although this insight would of course have been impossible without such prior study!), and leads to questions of what can be done with networks of neurons, and what naturally emerges. The ability to maintain a continuum of fixed points is not generic and requires tuning parameters, so how does it emerge? The suggested answer is through feedback, which provides a signal for how well the network is doing, and thus allows for further tuning of parameters towards the desired behaviour, which is illustrated by a surprising and elegant experiment involving confused goldfish!

Section VII broadly discusses biological measurement processes (like vision or chemical sensing) and their nearness to optimality. If operation near biological limits is the rule, then this provides a guide to selection of parameters - choose the parameter set which is as close as the system allows to the physical limits imposed upon it. Again neuronal networks provide a good example here, for instance the ability of flies to avoid being swatted. The combination of their fast movement speed and low-resolution eyes leads to a requirement for the fly brain to carry out near-exact motion estimation. This is a nice idea - it arises from a clear biological principle but can be formalised precisely in terms of information-theoretic concepts. It also ties in nicely to the principle in the previous section; information generates feedback, which helps with the precise tuning of a system to achieve the desired effect. From a thermodynamic perspective, however, more accurate measurement can often be carried out at an increasing energy cost (for example the Berg and Purcell work on cellular sensing). This implies a tradeoff: the organism can balance expenditure with increased information gain and it's not clear where this balance should be set. In non-sensing organs it's not clear what design principles could be at work - although information makes a lot of these ideas easy to quantify, it also limits their range of application.

Section VIII talks mainly about statistical physics-type models involving emergent behaviour from a large number of relatively simple agents. The main example here is the highly successful model of flocking in birds. This was revolutionised by the ability to take large-scale measurements of many (>1,000) individual birds simultaneously, which was used to build a maximum entropy flocking model. Local interactions are present, however long-range correlations emerge in the flock through Goldstone modes (in the case of direction) and tuning to a critical point in velocity. Further experiments show that criticality, rather than symmetry breaking, is the most common mechanism for the emergence of long-range correlations.

Although Section VIII is claimed to be a separate guiding principle earlier in the paper, the re-emergence of criticality as a central player suggests that really this is more an exploration of the consequences of the principle discussed in Section VII. Because of this it seems that there's really one principle at work here - emergent behaviour through tuning to criticality where the tuning occurs due to the flow of information in the system, which we expect to be as close to physical limits as possible. This seems like a fruitful avenue for the creation of characteristic models which capture these ideas together while remaining as simple as possible, and I'd be very interested to learn of any such models currently in existence.

Sunday, 13 September 2015

Shameless self-promotion: Engineering self-assembly pathways

Over the last two weeks we've finally managed to get our paper on self-assembly pathways published in Nature ("we" includes the Turberfield and Kwiatkowska groups in Oxford). In the paper, we look at a type of nanostructure called DNA origami.  DNA origami is a wonderfully versatile approach to create structures with a typical size of a few hundred nanometres, but with the possibility of adding details within a precision of a few nanometres. For real "wow-factor", just type "DNA origami" into Google image search.

The principle is remarkably simple. We start with one long DNA strand (see the pink and green object in the above picture) - this is known as the "scaffold". We then introduce a number of much shorter strands called "staples" (blue in the above picture). These staples are designed to stick to two (or more) specific parts of the scaffold very strongly, bringing them together and folding the scaffold into a complicated shape.

A number of people have made incredibly sophisticated structures, in both 2d and 3d, with this technique. However, we were more interested in understanding how the structures formed, and whether we could design them to form more efficiently. We therefore started with a simple rectangular design. Our unusual step was then to double the scaffold, so that it contained two identical halves (pink and green sections in the picture above). This meant that each staple could stick in a number of configurations, because each binding site appeared twice on the double scaffold.

Despite this, most scaffolds folded into one of a small number of distinct shapes, which looked like two rectangles lying side-by-side: example microscope images are shown in (b) and (c) above. We saw that certain shapes were more favourable than others. More interestingly, we were able to change which shapes were most common by making small adjustments to the staple strands. We were able to predict the changes using a theoretical model that we discuss in an even newer paper. The success of our model gives us some hope that we might be able to design origami rationally to improve the reliability of self-assembly.

From a general perspective of  using molecules to achieve complex tasks, the most interesting thing is that we were able to manipulate self-assembly outcomes by interfering with the folding pathway, rather than the stability of the final structures. The obvious way to force a system to assemble into a certain shape is to design your system so that your chosen structure has by far the lowest energy (technically, free energy) of all configurations. However, the alterations we made to staples should have had almost no effect on the relative energies of the different configurations. Instead, our modifications changed the order in which staples attached to the scaffold - and this order is crucial, as the early staples shape the scaffold, determining which of the possible structures will eventually form.




Saturday, 8 August 2015

Whaling on the nanoscale: Molecular harpoons

[Image of Type VI secretion system from the homepage of the Jensen Lab]

Greetings from Virginia, where I am currently at the Q-bio conference. I've seen plenty of great talks, and one audience member in the front row that slept the whole way through mine (he probably got up early to listen to the Ashes as well).

One talk by David Bruce Borenstein from the Wingreen group in Princeton brought the "Type VI secretion system" to my attention (for a more scientific discussion, see this review). This is a rather dry name for a pretty amazing bacterial weapon. Certain types of bacteria (in fact, quite a lot of them) have a harpoon concealed within their cell membrane that can be thrust into neighbouring cells, allowing delivery of toxic biochemicals. Bacteria use this weapon against each other and more complex organisms such as humans. It is similar to (and components may even have been directly stolen from) mechanisms by which some viruses inject genetic material into hosts.

For me, the interesting thing about this device is the ability to convert chemical changes of proteins within the cell into the rapid motion of this harpoon. For example, how exactly is chemical fuel involved, and how much fuel the cell must use to achieve a certain force? How much of an advantage does this active puncturing give?

Tuesday, 23 June 2015

Dissipative self-assembly

I'm currently at a conference on Engineering of Chemical Complexity hosted by the TUM in Munich. Today we had an interesting talk from Thomas Hermans on ``dissipative self-assembly". To understand this concept, we need to think about steady states.

At a first glance, many of the things around us don't change over time -- they are in steady states. For example, the hotel I'm currently sitting in looks pretty much the same as it did five minutes ago. A fundamental principle of physics is that isolated systems tend to relax towards an "equilibrium state" (this is essentially the famous second law of thermodynamics). Equilibrium is an example of a steady state, because when a system reaches equilibrium, it has nowhere else to go. It isn't, however, the only example. In fact, most objects that we see in steady states are actually stuck in "kinetic traps", including the hotel. Really, the equilibrium state of the materials that make up this hotel wouldn't look very welcoming! For a start, everything would be much closer to the ground. If we wait long enough, of course, the hotel would fall down. but the rate at which this happens is very slow and it seems apparently "trapped" in a hotel-like steady-state to the casual observer.

Dr. Hermans wants us to consider a third, fundamentally distinct type of "dissipative" steady state. In kinetically trapped or equilibrium steady states, the system maintains itself - you don't need to supply anything to keep it in the steady state. However, think about a human body which is, roughly speaking, in a steady state; it needs to be constantly supplied with food to stay this way. This situation is typical of many biological systems - they are out of equilibrium, but not kinetically trapped, and rapidly relax towards equilibrium unless they are fed with fuel in some form. Feeding them with  fuel keeps them in a "dissipative" steady state, which is called dissipative because fuel gets used up in the process.

Dr. Hermans is looking to design artificial biochemical assemblies that exist in such dissipative steady states. Why might this be worthwhile? Hallmarks of biological systems include their flexibility, repairability and adaptability, features that are probably much more natural in dissipative assemblies in which there is a constant turn-over of material. As yet, the results seem preliminary and I can't find any publications - although a discussion of the principle of dissipative self-assembly can be found in this article, "Droplets out of equilibrium".

Tuesday, 16 June 2015

Using a single molecular reaction to perform a complex calculation

Much of the research in molecular computing is based on making molecules imitate the most simple logical operations that underlie conventional digital electronics. For example, we can design molecular systems that mimic "AND", "OR"  and "NOT" gates. These gates can then be joined together to perform more complex tasks (eg. here). But at some level, physical systems naturally perform complex calculations. For example, the energy levels of a hydrogen atom are predicted theoretically through a number of complex equations, and in principle you could infer the answers to these equations by measuring the energies.

Of course, there are a number of issues with this. You need to be confident about your theory that relates measurements to equations, and you also need to be confident about the measurements themselves. But perhaps the most important caveat is that the equations that you can solve by performing measurements are often only interesting in predicting the outcome of the measurements themselves; the whole thing becomes rather incestuous and not obviously useful.

In a recent paper (here), Manoj Gopalkrishnan shows how the the complexity of a single chemical reaction can be harnessed to perform a useful computation. The paper is quite involved, but the essence is the following. Let's say I have the chemical species, X1, X2 and X3 that can interconvert by the reaction X1 + X3 ⇌ 2 X2, with both left and right sides being equally stable (such a system could, in principle, be created from DNA). If I start with a certain initial concentration of molecules, x1, x2 and x3, the system will evolve to reach some equilibrium in the long time limit, x'1, x'2 and x'3. This equilibrium represents a maximization of the system's entropy, subject to the constraint that (x1,x2,x3) can only be changed via the X1 + X3 ⇌ 2 X2 reaction.

Dr Gopalkrishnan shows that the constrained optimization performed by the chemical reaction solves a problem of wider interest. Let's say we have a random variable that can take three distinct values y1, y2 and y3 (in essence, a three-sided die). This die might be biased, with the probabilities of seeing each side not equal to 1/3. Further, we might have a reason to think that the probabilities are constrained by underlying physics, so that only certain combinations of probabilities p(y1), p(y2) and p(y3) are possible. So we know something about our die, but not the specifics. If someone rolls the die several times and presents us with the results, can we estimate the most likely values of p(y1), p(y2) and p(y3)?

Dr Gopalkrishnan shows that, for a certain very general type of constraint on probabilities (a "log-linear model"), the procedure for finding this "maximum likelihood estimate" of p(y1), p(y2) and p(y3) is exactly identical to that performed by an appropriate chemical system in maximizing its entropy under constrained particle numbers. Different log-linear constraints on p(y1), p(y2) and p(y3) translate into different choices of the reaction, which doesn't have to be X1 + X3 ⇌ 2 X2. Given the appropriate reaction, all we have to do is feed in the data (the results of the die rolls) as the initial conditions (x1,x2,x3), and the eventual steady state (x'1,x'2,x'3) gives us our estimate of (p(y1),p(y2),p(y3)).

In principle, this argument generalizes to much more complex systems than the illustration provided here, and the principle of maximum likelihood estimation isn't only applicable to biased dice. Of course, this is a long way from a physical implementation, and even further from an actual useful device, but it does illustrate the potential of harnessing the complexity of physical systems rather than trying to force them to approximate digital electronics. Going further in this direction, the chemical system will fluctuate about its steady-state average (x'1,x'2,x'3); it may be that these fluctuations can be directly related to our uncertainty in our estimate of (p(y1),p(y2),p(y3)).*


*In detail, the distribution of (x1,x2,x3) in the steady state may be related to the posterior probability of (p(y1),p(y2),p(y3)) given a flat prior and the data.