I know this blog is a little late to the game. In fact, I started writing it while it was still relevant... but then I got distracted. Such is life.
Background
You probably already know (or vaguely remember) what the Ice Bucket Challenge is. Basically if you're nominated you have to dump a bucket of ice water over your head, or else you have to donate money to some charity - most commonly ALSA, the Amyotrophic Lateral Sclerosis Association.
You then nominate 3 more people, who have 24 hours to do the same thing. In some variants, you also make a small donation even if you dump the ice water over your head, or else you make a bigger donation if you don't (some specify $100).
Anyway, what I was interested in is how the challenge spread - this was a 'viral' campaign in a very true sense. So why not try to model how the campaign spread as we would model the spread of a virus?
The Viral Model
I've written about this sort of thing several times before. The basic idea is this - the population is divided into three groups:
Susceptible (S) - The population that hasn't been exposed to a disease/virus, but is susceptible to infection
Infectious (I) - The people who have been exposed, and can infect other people
Recovered (R) - The people who have been infected, but have recovered. It's generally assumed that these people can't be re-infected. (But there are variations).
For the ice bucket challenge, we can look at a direct analogy as
S - Those who haven't been nominated
I - Those who have been nominated and are taking the challenge (so can nominate other people)
R - Those who have completed or those who declined the challenge (and can't nominate anyone else)
We can draw the model as a flowchart, showing how people move between the different groups,
Here, we've divided nominees into those who accept the challenge (S->I) and those who don't accept (S->R). So, now we can describe the model as a system of differential equations,
Where the factors (α, β, γ) describe what proportions of each group move to where.
The most interesting part is the 'S*I' terms in the first two equations - this basically says that the number of newly infected people is proportional to the number of susceptible people AND the number of infectious people.
This makes the system self-limiting, meaning that the number of infected people can't just grow to infinity - since that's not what we observe. Instead, what we see is an increase to some maximum, followed by a steady drop-off. For example,
We don't have hard data on how many people did the challenge over time, but we can get an idea of what happened by looking at the YouTube search numbers for the phrase 'ice bucket challenge' (above, via Google Trends).
I mean, it seems reasonable to assume some loose correlation between the number of people doing the challenge, the number of videos of people doing the challenge, and the number of searches for those videos. Searches peaked on August 21st, in case you were wondering.
A More Discrete Model
The downside to this 'S*I' term is that it makes the equations non-linear, meaning they can't be solved analytically - that is, you can't come up with an 'exact' equation for the number of people doing the challenge on a given day, for example.
But we can do a numerical simulation, instead.
In fact, this is a more reasonable way of looking at the system, since we're interested in a discrete time step of one day - i.e. the 24 hours nominees have to complete the challenge. For the simulation, we're also going to rounded the numbers of people in each group (after each time step) to whole numbers, since you can't (or shouldn't) divide a person into fractions.
Now, we need to define some parameters. First of all we need to define our start populations. We'll call the initial susceptible population S0 - this could be, for example, the Earth's population (which was around 7.16bn when I ran my simulations). We'll assume that one person is infected as a starting point - the originator of the challenge (I0 = 1). And we'll assume that no-one else has done the challenge at the start, so R0 = 0.
For the constants (α,β,γ), the easiest one to define is γ - the 'recovery' rate. We assume that after the 24 hour challenge period nominees are no longer 'infectious', therefore γ = 1.
For α and β, we have the total rate of infection/nomination defined as (α+β). We're given that each challenge completer gets to nominate three new people, so we can define (α+β) such that in the first step (when there's only one infectious person) we have (α+β)*S0*I0 = (α+β)*S0 = 3. Therefore (α+β) = 3/S0.
Now, α and β are related to the proportions of nominees that accept and decline the challenge, respectively. So we can redefine the constants as α = a/S0 and β = b/S0, such that (a+b) = 3, or alternatively α = a/S0 and β = (3-a)/S0. So now we can look at 'a' as the average number of nominees who accept the challenge.
So to tie it all together, we have the system of (difference) equations,
And it's pretty straightforward to write some code that'll run through those equations.
So we've got our model, what now?
Having a model is all well and good, but why bother? Well, now we can start asking questions. For example, how fast does the campaign spread? How long will it take for the challenge to die out, and how many people will have taken part by that point? And what happens when we change the number of people who accept the challenge?
First of all, if we run the simulation (with a = 2.5) and plot I(n) - the number of people doing the challenge on any given day - we get something like this
[where I(n) has been normalised so that the maximum is 100, as Google Trends does].
This is about what we expected - a rise and a fall. Though it's worth noting the shape is a little different from the YouTube searches above; the simulation drops off quickly, while the search numbers have a longer tail. This could be because, even after the challenges are done, there's a latent interest in (re)watching the videos. Or it could just be that this model is not entirely accurate...
So what happens when we vary 'a'?
We can start by assuming everyone accepts the challenge (a = 3). In this case, eventually everyone in the world does the challenge - specifically within 21 days of the first challenge. But that isn't very realistic.
So let's say that on average 1.5 or 2 of those challenged accept. In these cases, the ice bucket phenomenon ends before everyone can be challenged. For a = 2, the challenge ends after 50 day, with ~13% of the population going unchallenged. On the other hand, for a = 1.5 the challenge takes significantly longer to end (over a year).
In fact it turns out that for any a < 1.6030165.. the challenge will take a significant amount of time to end - the number of 'infectious' people will eventually reach 1, and stay there until S <= S0/(2a) (remember we're rounding each group to the nearest whole number). For a <= 0.5 the challenge ends almost immediately.
Now, we know that for a = 3, the population S goes to zero (everyone is challenged), whereas for a = 2 the challenge stops before S can go to zero. So we have the question - what is the smallest value of 'a' for which S goes to zero? If you do a bit of interpolating, you can figure out that this critical value comes out at around 2.7405349.. At this value of 'a', the challenge ends after 23 days. In other words, if on average 2.74.. (~91%) of the people nominated accept the challenge, then after 23 days everyone in the world will have been challenged.
From what I've seen myself, it seems like nearer 1 in 3 people accept the challenge on average. So, as viral as the campaign was, it was never going to take over the world.
If you play around a bit, you'll find that the critical values of 'a' are dependent on the initial population (S0). But, as far as I can tell, it's not possible to derive these critical values analytically. (Answers in the comments if you can prove otherwise).
Why wait to be nominated..?
At this point, you're probably thinking this isn't a very realistic model. And you'd be right. So lets make it a bit more complicated.
In particular, lets add spontaneous participation - that is, people who aren't directly nominated, but who see all the other people doing the challenge, and decide they want to take part too.
We'll assume that this participation is proportional to those who have already done the challenge (I and R). So to start with, we need to separate the 'recovered' group into those who actually did the challenge (R), and those who declined (D).
The new model looks something like this,
With difference equations,
In the flowchart, we've introduced this new factor 'δ', the 'inspiration' rate. If we re-define it, like we did for the infection rate, as d/S0, then 'd' can be loosely interpreted as the average number of people inspired to take part by each person who's already done the challenge.
Now we can look at how this 'd' factor affects the things we looked at before - how long it takes for the challenge to end, etc.
Let's start by assuming that the average number of nominees accepting the challenge, 'a', is 1 (out of 3). What value of 'd' do we need for S to go to zero? Do a little interpolating again and you get d = 0.8668653 - that is, if each person who pours a bucket of water over their head inspires (on average) 0.867.. people to do the same, then everyone in the world will participate, within 26 days of the challenge starting. For a = 2, we need d = 0.4046021 for S goes to 0. And so on...
What's a plausible inspiration rate? For a normal person, probably zero, while for a celebrity... I don't know. But on average 'd' is probably very close to zero. I mean, we know for a fact that significantly less than the entire population of Earth has been nominated/taken part in the challenge.
If you plot I(n) for a = 1 and d = 0.1 you get something like this,
In this case, you have that same rise and fall - but this time, you have a longer tail, like we see in the YouTube searches. Is this proof that this iteration of the model is more accurate? Maybe. But as I pointed out before, the YouTube searches don't necessarily accurately represent the number of people taking the challenge over time.
Social Pressure
When a friend does a thing for charity, then publicly calls you out to take part, there's a certain amount of social pressure to comply. I mean, if a friend dumped water over their head (just because), then asked you to do the same thing, probably you'd look at them like they were a crazy person. Anyway, that kind of social pressure is implicitly included in the 'infection' rate - more social pressure, bigger 'a'.
Instead, the sort of social pressure I'm talking about here is the kind that goes: "I should accept the challenge because so many other people have already done it". Or alternatively, "it's okay for me to accept the challenge, since so many other people have already done it".
In other words, the more people accept and complete the
challenge, the more likely a nominated person is to accept too. Mathematically, we can introduce this with a term in 'S*I*R'. Or alternatively, we can keep the term as 'a*S*I', but make the factor 'a' a function of R -> a(R).
Anyway, if you're interested, you can try investigating that yourself. Or try adapting the model in some other ways. But beyond a point you can end up complicating a model more than improving it.
A Network Theory Approach to Nominating
So I was eventually nominated for the challenge. But I'm a wimp, so I declined to dump ice water over my head, instead making a donation to the Motor Neuron Disease Association (the UK equivalent of ALSA).
For my nominations, I wanted to try and maximise spread. So I nominated the 3 of the people in my Twitter network who are the most active and well connected, and who I thought would be up for accepting the challenge. Plus, as a secondary effect, I figured they might nominate other people in my Twitter network - the network theory equivalent of wishing for more wishes.
In the end, one ignored the nomination, one acknowledged but didn't accept, and one accepted (in the form of a donation) but didn't nominate anyone else. So I guess that theory didn't quite pan out. But I did at least encourage more charitable giving.
So Yeah
If we had real world data we could maybe test the accuracy of these models. But even without, we can get a sense of how the challenge behaves - for example, we find that there's a critical 'challenge acceptance ratio' that determines whether the viral campaign will go 'pandemic'.
In theory, you could apply this sort of model to any viral campaign, or just anything that spreads 'virally'. The nice thing about the Ice Bucket Challenge in particular, though, is that it has well defined rules for how the challenge spreads from one person to the next.
So, yeah..
Oatzy.
[I need an editor, my pronouns are all over the place.]
Showing posts with label sociology. Show all posts
Showing posts with label sociology. Show all posts
Thursday, September 18, 2014
Sunday, December 22, 2013
Modelling the Survival of Chocolates
So it's been about a year since I last blogged. I've been kinda busy with university stuff, and growing my facial hair into something approaching a beard. I might write a blog about it at some point.
Anyway, I've got a little free time now - at least until after Christmas - so here we are. Let's waste some time...
"This is a real paper, published in a proper journal"
So the other day, my dad hands me this print out of an academic paper, "this seems like your sort of thing". The paper is called
"The survival time of chocolates on hospital wards: covert observational study" [1]
The basic set-up is this: staff at various hospitals put out boxes of Roses and Quality Street chocolates on wards, and (covertly) recorded every time one was taken. From this, they analysed the 'survival' of the chocolates, according to Kaplan-Meier survival analysis.
They found that the rate of chocolate consumption fit an exponential decay, with the median survival time for a chocolate being 51 minutes, and the survival half life ("time taken for 50% of the chocolates to be eaten") being 99 minutes. They also found that Roses chocolates were preferred over Quality Street.
"Damn it, man, I'm a physicist, not a doctor!"
So I'm reading this paper, and I'm thinking "I reckon I could model this". So I had a crack at it.
Basically, what we've got here is an agent-based modelling problem. Individuals encounter a box of chocolates at random intervals. Each box has a variety of different types of chocolates. So each individual will look through the box for any chocolate(s) they like the look of, and help themselves.
We can make a few simplifications to this. First of all, we work in fixed, discrete time intervals. In each time interval, some number of people will encounter the box of chocolates - one person per interval in the simplest model, and some random number of people in a more complex model.
Now, different people have different preferences. We could generate for each person a random set of preferences, but for simplicity we'll assume some 'average preference probability'. In other words, imagine there are equal numbers of each type of chocolates - what is the probability that any given person will pick type 1, type 2, etc. ?
So, in the model, each person will pick a random chocolate based on this preference probability distribution. This accounts for the fact that some types of chocolate (strawberry) are naturally more popular than others (toffee penny).
Here's Where it Gets Maths-y
We have a box of N chocolates total. There are m different types of chocolate, with each type i=1,..,m having a count of ni such that
Each type also has a 'preference probability' pi with 0 <= pi <= 1 and
In each time step dt an individual randomly chooses a type of chocolate, i in {1,...,m}, based on the preference distribution, and one chocolate of that type is removed from the box according to
This process is repeated until N(t) = 0
In the more complex model, some random number of individuals (following a Poisson distribution) each choose chocolates in each given time step. This is also functionally equivalent to a single individual choosing multiple chocolates in a single given time step.
What Differential Does It Make?
In the simplest model, one person chooses a single chocolate in each time step. This is nice, because it means we can express the model as a set of differential equations, which can be solved exactly.
For each type of chocolate we have the rate of consumption given by
which has standard (exponential decay) solution
And since the total count N is just the sum of the individual counts ni we get the full solution
In other words, it's just the sum of several exponential decays (with different decay rates) - making it, roughly, an exponential decay itself.
In particular, in the special case where all the chocolates are equally preferred such that pi = p for all i=1,..,m we get the solution
And all that is loosely what was observed on the hospital wards.
But these 'exact' solutions don't capture some of the subtleties of the real life set-up. For example, the toffee pennies might be completely ignored until they are the only option left; whereas, the exact solution assumes that (fractions of) all of the types of chocolate are being taken all the time, just at different rates.
"We did not seek ethical approval for this study..."
As an improvement on the analytical solution, we can create a computer simulated model, that captures more of the randomness of real life (code below). For the smplest model, the results look something like this
Notice that in this version of the model, the number of chocolates decays more slowly, and has some of the granularity seen in the observed data plot.
Plus, the randomness means the specific behaviour varies from run to run - for example, the orange line decays relatively quickly, while the blue has a long stretch when nothing is being taken.
Interestingly, this model seems to have a linear decrease for the first ~20-30 time steps in every sample run. This is probably because, for those first few steps, everyone can get what they want. But eventually, certain types run out, and the rate of consumption decays away.
Why Do Chocoholics Come in Threes?
I wrote about the Poisson distribution before, back when I was modelling cafes. Basically, while we might see an average of one person per time interval, we don't expect to see them so evenly spaced out. Rather, people will arrive in 'bursts' of varying size. This behaviour is described by the Poisson distribution.
So I wrote my simulation in Python, using the numpy library for the Poisson distributed random numbers. The results look something like this
In this case, the initial decrease is less linear than in the simple model, but the overall decay times are roughly the same (except for the blue line above). This is what we'd expect since, in the above, I chose the average number of people per interval to be 1.
So yeah. Try the code, have a play with the parameters, see what happens. The output is the total number of chocolates remaining after each time step.
One thing to look at would be how changing the probabilities affects the rate of decay - for example, does the number of chocolates decay faster when all the chocolates are equally popular, or when some are much more popular than others. I suspect the former is the case.
It might also be worth adding an option where a person can pick again if their prefered chocolate isn't avaiable.
And the IgNobel Prize for Medicine goes to...
So that's that. And I can definitely see the original paper being nominated for an IgNobel, at the very least.
As for me, I have some gift wrapping I really need to get to...
Merry Christmas!
Citation
[1] P R Gajendragadkar, et al. "The survival time of chocolates on hospital wards: covert observational study" BMJ 2013;347:f7198 (Published 14 December 2013)
Oatzy.
[Next job, modelling Her Majesty's nuts...]
Labels:
analytical,
chocolate,
everyday maths,
maths,
modelling,
problem solving,
programming,
simulation,
sociology
Tuesday, May 29, 2012
Sexuality and Population Dynamics
So I was reading this response to a Daily Mail opinion piece about those 'gay cure' bus adverts, and I see this quote (from the Mail article)
First of all, that 1% values. The actual number is a tricky figure to pin down. The UK Office of National Statistics puts the figure at 1.5% gay or bisexual; the linked response lists 3 studies which put the value around 7% (for the UK), and various other studies suggest number in the modern West is somewhere between 2-13%.
Secondly, just because something is (statistically) "a departure from the norm", doesn't mean it's bad - for example, ~10% of people are left-handed, ~2% of people have red hair, etc. (Historical beliefs notwithstanding.)
But all that's tangential. What I'm interested in is the second part - would a majority gay population spell the end of humanity?
[NB/ for the remainder of the blog, I shall be using 'gay' colloquially, as shorthand for all of LGBTQUA, etc.]
The Simple Model
It's an interesting population dynamics question, and it's easy to model if we make some assumptions and simplifications:
1) We assume that homosexuals don't reproduce at all (either directly or indirectly). This isn't true - but it seems to be a common belief.
2) We assume that the rest of the population reproduces with some fixed rate (b=0.02059), which doesn't change from year to year. Similarly, we assumes a fixed death rate (d=0.00812), which affects all members of the population equally.
3) We assume that sexuality is binary, definite and static, with a fixed incidence of homosexuality, g = 0.07 - that is, of all children born 7% will be gay.
So the model we're using is preposterously simplified - effectively,
Now, there are two ways to interpret the article's statement in the context of our model -
1) "if 99% of the (initial) population were gay" assuming the 7% incidence
The outcome looks like this
In this case, the population drops at first, but then starts to grow again, exponentially - certainly not the end of the human race. Also notable is that the proportion of the population that is gay stabilises at 7% (as expected).
2) "if 99% of people were born gay"
Now, the outcome looks like this
Exponential decay - so, in this case, humanity would die out. But if 99% of the population isn't reproducing, what do you expect?
Here are the phase portraits of the above two cases, for a slightly different view of what's happening (gay population along x, straight population along y)
Again, for g=0.07 the population tends to infinity; for g=0.99 the population tends to zero. Note that this behaviour is (ultimately) the same for all initial population ratios.
What I Learned in MAS271
We don't actually need numerical simulations to study the behaviour of this (or similar) systems - the model can be solved analytically.
First of all, we generalise the system. We start with a pair of coupled, first-order differential equations describing the behaviour of the two populations
Now, these two populations are mutually exclusive subgroups of the total population - that is, total population z = x + y. So with a little mathematical trickery, we can combine these two into a single equation, describing the behaviour of the total population
This is a second-order, linear differential equation, so has the general solution of the form
Where lambda values are given by
And A-coefficients are constants, which are determined by the system's initial conditions. If the initial total population is z0, and the inital size of the x population is x0, then we get
In other words, we now have an (analytical) equation which completely describes how the population will behave for any given set of variable.
Incidentally, the equations for x and y have the same form as z. The only difference is the A-coefficients.
Boom or Bust
What we want to do is be able to determine the long term behaviour of the population based on the given set of variables - that is, will the population grow, or die out?
For the coupled system above, there is only one stationary point, at (x,y) = (0,0) - i.e. when the total population is zero. What this means is, if the population is zero, it will stay zero. If it has any other value, then the population size will change over time.
In terms of the stability of the system, if (0,0) is stable, the population will tend towards zero - i.e. will die out. If (0,0) is unstable, the population will tend to infinity - i.e. grow indefinitely.
As it turns out, the lambda values above are the eigenvalues of the coupled system, meaning their values determine the stability of the system:
i) if both lambdas are negative, the system is stable and the population will die out;
ii) if one or both of the lambdas are positive, the system is unstable and the population will grow to infinity.
Now, if we take the limit of z(t) for t tends to infinity, we can simplify the population equation
What this tells us is, if lambda1 is less than one the population will die out. If lambda1 is greater than one, the population will grow.
And we can use this to determine the conditions for instability - the conditions under which the population will grow. In this case the condition is lambda1 > 0. But this can be simplified down to
There's one last trick we can do - determine what proportion of the total population will be in the sub-population x (e.g. what what fraction of the population will be gay) as t tends to infinity.
For this, we determine the equation for population x (in a similar fashion to z), and take advantage of the limits to get the equation
Note that this equation is time-independant, i.e. the proportion will tend to some constant value.
Back to Basics
Which gives the population equation
Which behaves like the graphs at the top. And this tends to
And gives
That is, the proportion of the population that will be gay stabilises at the pre-defined 'birth' ratio - which is what we saw from the numerical model. Note also that this value is independent of initial conditions and doesn't vary over time.
Simple Instability
Finally, for the simple model, the condition for instability (growth) is given by
For the numerical values chosen above, this gives the condition g < 0.606 - that is, if less than 61% of the population is 'born' gay, then the population won't die out. And 61% is a pretty high upper limit - a majority, in fact.
Similarly, we can use this inequality to find under what conditions g=0.99 is viable. In this case, it requires a birth rate 100 times the death rate.
So, if the death rate is 0.008 (i.e. 8 death for every 1,000 people), then it requires a birth rate of 0.8 (800 birth per 1,000 people). Now, for each 1,000 people, 99% is gay, meaning 10 straight people per 1,000. Of those, about half is female (5). In other words, each woman would have to give birth to 160 babies per year!
Clearly this isn't feasible. But for comparisons sake, consider bees or ants - these species are eusocial, meaning populations made of one fertile queen, and hundreds of (mostly) sterile drones. In this case, the high birth-to-death ratio is more reasonable, and the predominantly non-reproducing population is still viable.
A More Complex Model
In the simple model we assumed that only the straight population reproduces, and that both populations die at the same rate. So we make the following modifications:
1) First of all, lets tweak the model so that the gay population has children as well, with some separate birth rate, b1.
2) Of those born to gay couples, some proportion g1 will be gay (NB/ studies show that the majority of children raised by same-sex parents grow up to be heterosexual).
3) And for completeness, lets say that the gay population has a different death rate as well, d1.
Here's what the model looks like
And here are the coupled equations
Now, all the previously derived general equations and such from earlier still apply. But in this case, the variables can't really be simplified, and writing them out in full would just be a mess.
So if we take, for example, b and d as before, and b1 = 0.00103, d1 = 0.01218, and g1 = 0.03, here are the new phase diagrams for g = 0.07 and 0.99
For this model, the phase portraits are basically the same as for the simple model - i.e. the effect of using the modified model over the simple model is negligible.
Complex Conditions
So, for this model, the condition for growth looks like this
The important thing to note here is, if the condition for the simple model (g<1-d/b) is satified, then the condition for the complex model is necessarilty satisfied.
In fact, if we set g to its maximum possible value 1 (i.e. all babies born to straight couple are gay), then the human race will still persist so long as g1 > d1/b1.
And what this means is, so long as the birth rate (b1) is greater than the death rate (d1), then every baby born could be gay, the human race still wouldn't die out.
However, given the low birth rate for gay couples (compared to death rate), this might not actually be satisfied IRL. In which case, an entirely gay population is still not sustainable.
For the b and d values already chosen, we can use the condition to plot how the upper limit of g varies with g1
So for these values, the upper limit of g can increase by no more than 3.4%. So, again, the extra complexity doesn't give a significantly different result compared to the simple model.
In fact, the additional variables only become significant when (b1/d1)g1 > 0.15 (for the values above); or in general, when
So, if we keep b and d as before, g=0.07, and re-set b1=b, d1=d, and g1=0.06, then (b1/d1)g1=0.152, and the phase portrait looks like this
Here, the upper limit on g has risen to 0.67. (cf: 0.61 for the simple model). And, if we set the initial condition x0=0.07*z0, then the limit of x/z = 0.0697; just less than g.
tl;dr - So, would homosexuality ever mean the end of the human race? - Under some circumstances, yes. But in most (real world) cases, no. And even if a majority of the population were LGBT, the human race still wouldn't necessarily die out.
BUT, as a final note - remember that these are very simplified models of population dynamics. So take the results with a grain of salt. The science of sexuality is much more complex and much less abstract than this.
Oatzy.
[Putting my degree to good use.]
“Since only about one percent of us are [gay], homosexuality is obviously a departure from the norm... Reversing that proportion would spell the end of the human race, which is clearly undesirable.”There are several things wrong with this statement - some of which the response, linked above, points out.
First of all, that 1% values. The actual number is a tricky figure to pin down. The UK Office of National Statistics puts the figure at 1.5% gay or bisexual; the linked response lists 3 studies which put the value around 7% (for the UK), and various other studies suggest number in the modern West is somewhere between 2-13%.
Secondly, just because something is (statistically) "a departure from the norm", doesn't mean it's bad - for example, ~10% of people are left-handed, ~2% of people have red hair, etc. (Historical beliefs notwithstanding.)
But all that's tangential. What I'm interested in is the second part - would a majority gay population spell the end of humanity?
[NB/ for the remainder of the blog, I shall be using 'gay' colloquially, as shorthand for all of LGBTQUA, etc.]
The Simple Model
It's an interesting population dynamics question, and it's easy to model if we make some assumptions and simplifications:
1) We assume that homosexuals don't reproduce at all (either directly or indirectly). This isn't true - but it seems to be a common belief.
2) We assume that the rest of the population reproduces with some fixed rate (b=0.02059), which doesn't change from year to year. Similarly, we assumes a fixed death rate (d=0.00812), which affects all members of the population equally.
3) We assume that sexuality is binary, definite and static, with a fixed incidence of homosexuality, g = 0.07 - that is, of all children born 7% will be gay.
So the model we're using is preposterously simplified - effectively,
new population = old population + births - deaths
And we can easily program this model and run simulations.Now, there are two ways to interpret the article's statement in the context of our model -
1) "if 99% of the (initial) population were gay" assuming the 7% incidence
The outcome looks like this
In this case, the population drops at first, but then starts to grow again, exponentially - certainly not the end of the human race. Also notable is that the proportion of the population that is gay stabilises at 7% (as expected).
2) "if 99% of people were born gay"
Now, the outcome looks like this
Exponential decay - so, in this case, humanity would die out. But if 99% of the population isn't reproducing, what do you expect?
Here are the phase portraits of the above two cases, for a slightly different view of what's happening (gay population along x, straight population along y)
Again, for g=0.07 the population tends to infinity; for g=0.99 the population tends to zero. Note that this behaviour is (ultimately) the same for all initial population ratios.
What I Learned in MAS271
We don't actually need numerical simulations to study the behaviour of this (or similar) systems - the model can be solved analytically.
First of all, we generalise the system. We start with a pair of coupled, first-order differential equations describing the behaviour of the two populations
Now, these two populations are mutually exclusive subgroups of the total population - that is, total population z = x + y. So with a little mathematical trickery, we can combine these two into a single equation, describing the behaviour of the total population
This is a second-order, linear differential equation, so has the general solution of the form
Where lambda values are given by
And A-coefficients are constants, which are determined by the system's initial conditions. If the initial total population is z0, and the inital size of the x population is x0, then we get
In other words, we now have an (analytical) equation which completely describes how the population will behave for any given set of variable.
Incidentally, the equations for x and y have the same form as z. The only difference is the A-coefficients.
Boom or Bust
What we want to do is be able to determine the long term behaviour of the population based on the given set of variables - that is, will the population grow, or die out?
For the coupled system above, there is only one stationary point, at (x,y) = (0,0) - i.e. when the total population is zero. What this means is, if the population is zero, it will stay zero. If it has any other value, then the population size will change over time.
In terms of the stability of the system, if (0,0) is stable, the population will tend towards zero - i.e. will die out. If (0,0) is unstable, the population will tend to infinity - i.e. grow indefinitely.
As it turns out, the lambda values above are the eigenvalues of the coupled system, meaning their values determine the stability of the system:
i) if both lambdas are negative, the system is stable and the population will die out;
ii) if one or both of the lambdas are positive, the system is unstable and the population will grow to infinity.
Now, if we take the limit of z(t) for t tends to infinity, we can simplify the population equation
What this tells us is, if lambda1 is less than one the population will die out. If lambda1 is greater than one, the population will grow.
And we can use this to determine the conditions for instability - the conditions under which the population will grow. In this case the condition is lambda1 > 0. But this can be simplified down to
There's one last trick we can do - determine what proportion of the total population will be in the sub-population x (e.g. what what fraction of the population will be gay) as t tends to infinity.
For this, we determine the equation for population x (in a similar fashion to z), and take advantage of the limits to get the equation
Note that this equation is time-independant, i.e. the proportion will tend to some constant value.
Back to Basics
Just as a sort of check, we'll go back to the basic model. For this, we have
Which gives lambda and A valuesWhich gives the population equation
Which behaves like the graphs at the top. And this tends to
And gives
That is, the proportion of the population that will be gay stabilises at the pre-defined 'birth' ratio - which is what we saw from the numerical model. Note also that this value is independent of initial conditions and doesn't vary over time.
Simple Instability
Finally, for the simple model, the condition for instability (growth) is given by
For the numerical values chosen above, this gives the condition g < 0.606 - that is, if less than 61% of the population is 'born' gay, then the population won't die out. And 61% is a pretty high upper limit - a majority, in fact.
Similarly, we can use this inequality to find under what conditions g=0.99 is viable. In this case, it requires a birth rate 100 times the death rate.
So, if the death rate is 0.008 (i.e. 8 death for every 1,000 people), then it requires a birth rate of 0.8 (800 birth per 1,000 people). Now, for each 1,000 people, 99% is gay, meaning 10 straight people per 1,000. Of those, about half is female (5). In other words, each woman would have to give birth to 160 babies per year!
Clearly this isn't feasible. But for comparisons sake, consider bees or ants - these species are eusocial, meaning populations made of one fertile queen, and hundreds of (mostly) sterile drones. In this case, the high birth-to-death ratio is more reasonable, and the predominantly non-reproducing population is still viable.
A More Complex Model
In the simple model we assumed that only the straight population reproduces, and that both populations die at the same rate. So we make the following modifications:
1) First of all, lets tweak the model so that the gay population has children as well, with some separate birth rate, b1.
2) Of those born to gay couples, some proportion g1 will be gay (NB/ studies show that the majority of children raised by same-sex parents grow up to be heterosexual).
3) And for completeness, lets say that the gay population has a different death rate as well, d1.
Here's what the model looks like
And here are the coupled equations
Now, all the previously derived general equations and such from earlier still apply. But in this case, the variables can't really be simplified, and writing them out in full would just be a mess.
So if we take, for example, b and d as before, and b1 = 0.00103, d1 = 0.01218, and g1 = 0.03, here are the new phase diagrams for g = 0.07 and 0.99
For this model, the phase portraits are basically the same as for the simple model - i.e. the effect of using the modified model over the simple model is negligible.
Complex Conditions
So, for this model, the condition for growth looks like this
The important thing to note here is, if the condition for the simple model (g<1-d/b) is satified, then the condition for the complex model is necessarilty satisfied.
In fact, if we set g to its maximum possible value 1 (i.e. all babies born to straight couple are gay), then the human race will still persist so long as g1 > d1/b1.
And what this means is, so long as the birth rate (b1) is greater than the death rate (d1), then every baby born could be gay, the human race still wouldn't die out.
However, given the low birth rate for gay couples (compared to death rate), this might not actually be satisfied IRL. In which case, an entirely gay population is still not sustainable.
For the b and d values already chosen, we can use the condition to plot how the upper limit of g varies with g1
So for these values, the upper limit of g can increase by no more than 3.4%. So, again, the extra complexity doesn't give a significantly different result compared to the simple model.
In fact, the additional variables only become significant when (b1/d1)g1 > 0.15 (for the values above); or in general, when
So, if we keep b and d as before, g=0.07, and re-set b1=b, d1=d, and g1=0.06, then (b1/d1)g1=0.152, and the phase portrait looks like this
Here, the upper limit on g has risen to 0.67. (cf: 0.61 for the simple model). And, if we set the initial condition x0=0.07*z0, then the limit of x/z = 0.0697; just less than g.
tl;dr - So, would homosexuality ever mean the end of the human race? - Under some circumstances, yes. But in most (real world) cases, no. And even if a majority of the population were LGBT, the human race still wouldn't necessarily die out.
BUT, as a final note - remember that these are very simplified models of population dynamics. So take the results with a grain of salt. The science of sexuality is much more complex and much less abstract than this.
Oatzy.
[Putting my degree to good use.]
Saturday, March 24, 2012
Some Twitter Infographics
I did some stuff like this before. And I figured, while I was updating my network graphs, why not update some of the other graphics?
And it helps that I worked out how to easily extract data from Twitter (see previous blog). The code is here. Again, rate limits apply.
Who Do I Follow?
This is one of the ones I did before - collect together the bios of the people I follow, then make a word cloud (using Wordle)
Basically, I follow a bunch of geeks and writers. Who like 'things'. So really, same as a year and a half ago.
I would point out though that 6 of the people I follow don't have bios, and about 7 just have lyrics.
Data here.
Who Tweets the Most?
These rates are worked out as (total tweets posted)/(total days online). Obviously, the actually post rate will vary over different time scales..
Bubble chart (made with ManyEyes) - bubbles sized by tweet rate (the numbers on some of the bubbles).
The graph below gives a better idea of relative rates, and 'rankings' (click to embiggen)
The blue line is actual values.
The orange is a logarithmic trend-line. It's a pretty good fit (R2=0.95); and, loosely speaking, it means ~70% of the tweets in my timeline come from ~30% of the people I follow. [cf: Pareto Principle]
You get similar log-shaped graphs when you split up the genders.
Full data here.
Chattiest Gender?
You can read all the explanation, caveats, etc. in the previous posts (here and here). I'm just going to go straight into the data.
I follow 27 men and 21 women (excluding celebrities, etc.). The stats are as follow:
Basically, the women tweet more on average, and their rates are more spread out than for the men. In fact, roughly three quarters of the men tweet less than half of the women. Also, there's one outlier in the female group.
This is similar to what we found last time; although the women's average and spread aren't quite as high (average: 13.79 vs 19.21), and the men's average has increased slightly (6.21 vs 5.29).
If you take the ratio of the averages, the women tweet 2.15 times as much as the men. But maybe I just follow particularly chatty women..
Here's treemap (ManyEyes), which should give you a better idea of the gender balance (boxes sized by tweet rate)
Specifically, the graphic above is 62.5% purple (female).
Data here.
Where in the World Are My Followers?
The site I used last time doesn't seem to exist anymore. So I'm using MapMyFollowers instead. As the name suggests, these are my followers, rather than just the people I follow. Nonetheless..
Mostly in the UK and the US. As you'd probably expect.
I will point out though, some of the locations are a little suspect. Some people haven't made their location available so aren't included, and others seem to be in countries they couldn't possibly be in. But it's the best we can do.
Here's a zoom in on the UK
What Do I Tweet?
Made with Wordle, with data from TweetStats.
Words are sized by how often I tweet them; and by extension, @usernames are sized by how often I tweet those people.
In fact, here are the people I 'mention' the most (TweetStats)
Couldn't get a good source on who @replies me. That was one of the things Twoolr used to do..
When Do I Tweet?
Twoolr used to be awesome for Twitter statistics. But sadly, when they left beta, they started charging. And their free service went to shit. Luckily, I found TweetStats. Weirdly, it doesn't need you to log-in or anything, but somehow it can pull data on (nearly) all your tweets - beyond the 3,200 limit. Strange.
Here's some more graphs
Basically, I tweet most on a Friday and Saturday, and at around 1-2pm.
And I've never tweeted at 5am. But that's probably because I'm always asleep at 5am
Except that one time I got really drunk. (SleepBot)
How Much Do I Tweet?
This is another one I used to go to Twoolr for. And, to be fair, I still could. But that only goes as far back as April '10, and its graphics aren't as clear. Here's TweetStats again
Like I said before, I didn't tweet much in my first year. In fact, I only posted 36 tweets in all of 2009.
Now, the one problem with TweetStats is that 5 month gap in 2010. Why is this significant? Well, I was definitely tweeting during that time. In fact, by my estimates, over those 5 months I posted 5,724 tweets (~37tweets/day). So those 5 months account for 43% of all my tweets.
See, the thing is, in 2010, I was out of university, single, and unemployed. I posted a total 8,823 tweets - 24tweets/day. Since I've been back at university, that number's dropped to 11tweets/day.
That lull in Summer 2011 was when I was spending all my time on Tumblr and watching classic Doctor Who. Incidentally, I haven't posted on Tumblr since the start of September '11. It's terribly addictive, you see. I wouldn't recommend it; unless you're addicted to Doctor Who and Sherlock, and have lots of time on your hands..
So yeah.
Oatzy.
[Self-indulgent statistics, and pretty illustrations.]
And it helps that I worked out how to easily extract data from Twitter (see previous blog). The code is here. Again, rate limits apply.
Who Do I Follow?
This is one of the ones I did before - collect together the bios of the people I follow, then make a word cloud (using Wordle)
Basically, I follow a bunch of geeks and writers. Who like 'things'. So really, same as a year and a half ago.
I would point out though that 6 of the people I follow don't have bios, and about 7 just have lyrics.
Data here.
Who Tweets the Most?
These rates are worked out as (total tweets posted)/(total days online). Obviously, the actually post rate will vary over different time scales..
Bubble chart (made with ManyEyes) - bubbles sized by tweet rate (the numbers on some of the bubbles).
The graph below gives a better idea of relative rates, and 'rankings' (click to embiggen)
The blue line is actual values.
The orange is a logarithmic trend-line. It's a pretty good fit (R2=0.95); and, loosely speaking, it means ~70% of the tweets in my timeline come from ~30% of the people I follow. [cf: Pareto Principle]
You get similar log-shaped graphs when you split up the genders.
Full data here.
Chattiest Gender?
You can read all the explanation, caveats, etc. in the previous posts (here and here). I'm just going to go straight into the data.
I follow 27 men and 21 women (excluding celebrities, etc.). The stats are as follow:
Men:For clarity, here's a boxplot (made in R)
Average = 6.21 tweets/day
Standard Deviation = 6.47
Women:
Average = 13.79 tweets/day
Standard Deviation = 14.27
Basically, the women tweet more on average, and their rates are more spread out than for the men. In fact, roughly three quarters of the men tweet less than half of the women. Also, there's one outlier in the female group.
This is similar to what we found last time; although the women's average and spread aren't quite as high (average: 13.79 vs 19.21), and the men's average has increased slightly (6.21 vs 5.29).
If you take the ratio of the averages, the women tweet 2.15 times as much as the men. But maybe I just follow particularly chatty women..
Here's treemap (ManyEyes), which should give you a better idea of the gender balance (boxes sized by tweet rate)
Specifically, the graphic above is 62.5% purple (female).
Data here.
Where in the World Are My Followers?
The site I used last time doesn't seem to exist anymore. So I'm using MapMyFollowers instead. As the name suggests, these are my followers, rather than just the people I follow. Nonetheless..
Mostly in the UK and the US. As you'd probably expect.
I will point out though, some of the locations are a little suspect. Some people haven't made their location available so aren't included, and others seem to be in countries they couldn't possibly be in. But it's the best we can do.
Here's a zoom in on the UK
What Do I Tweet?
Made with Wordle, with data from TweetStats.
Words are sized by how often I tweet them; and by extension, @usernames are sized by how often I tweet those people.
In fact, here are the people I 'mention' the most (TweetStats)
Couldn't get a good source on who @replies me. That was one of the things Twoolr used to do..
When Do I Tweet?
Twoolr used to be awesome for Twitter statistics. But sadly, when they left beta, they started charging. And their free service went to shit. Luckily, I found TweetStats. Weirdly, it doesn't need you to log-in or anything, but somehow it can pull data on (nearly) all your tweets - beyond the 3,200 limit. Strange.
Here's some more graphs
Basically, I tweet most on a Friday and Saturday, and at around 1-2pm.
And I've never tweeted at 5am. But that's probably because I'm always asleep at 5am
Except that one time I got really drunk. (SleepBot)
How Much Do I Tweet?
This is another one I used to go to Twoolr for. And, to be fair, I still could. But that only goes as far back as April '10, and its graphics aren't as clear. Here's TweetStats again
Like I said before, I didn't tweet much in my first year. In fact, I only posted 36 tweets in all of 2009.
Now, the one problem with TweetStats is that 5 month gap in 2010. Why is this significant? Well, I was definitely tweeting during that time. In fact, by my estimates, over those 5 months I posted 5,724 tweets (~37tweets/day). So those 5 months account for 43% of all my tweets.
See, the thing is, in 2010, I was out of university, single, and unemployed. I posted a total 8,823 tweets - 24tweets/day. Since I've been back at university, that number's dropped to 11tweets/day.
That lull in Summer 2011 was when I was spending all my time on Tumblr and watching classic Doctor Who. Incidentally, I haven't posted on Tumblr since the start of September '11. It's terribly addictive, you see. I wouldn't recommend it; unless you're addicted to Doctor Who and Sherlock, and have lots of time on your hands..
So yeah.
Oatzy.
[Self-indulgent statistics, and pretty illustrations.]
Labels:
data analysis,
data tracking,
gender,
graph,
infographic,
nerd,
OCD,
programming,
python,
sleep,
sociology,
statistics,
tumblr,
tweets,
twitter
Subscribe to:
Posts (Atom)























































