Showing posts with label viruses. Show all posts
Showing posts with label viruses. Show all posts

Thursday, September 18, 2014

#IceBucketChallenge - A Viral Campaign

I know this blog is a little late to the game. In fact, I started writing it while it was still relevant... but then I got distracted. Such is life.


Background

You probably already know (or vaguely remember) what the Ice Bucket Challenge is. Basically if you're nominated you have to dump a bucket of ice water over your head, or else you have to donate money to some charity - most commonly ALSA, the Amyotrophic Lateral Sclerosis Association.

You then nominate 3 more people, who have 24 hours to do the same thing. In some variants, you also make a small donation even if you dump the ice water over your head, or else you make a bigger donation if you don't (some specify $100).

Anyway, what I was interested in is how the challenge spread - this was a 'viral' campaign in a very true sense. So why not try to model how the campaign spread as we would model the spread of a virus?


The Viral Model

I've written about this sort of thing several times before. The basic idea is this - the population is divided into three groups:

Susceptible (S) - The population that hasn't been exposed to a disease/virus, but is susceptible to infection
Infectious (I) - The people who have been exposed, and can infect other people
Recovered (R) - The people who have been infected, but have recovered. It's generally assumed that these people can't be re-infected. (But there are variations).

For the ice bucket challenge, we can look at a direct analogy as

S - Those who haven't been nominated
I - Those who have been nominated and are taking the challenge (so can nominate other people)
R - Those who have completed or those who declined the challenge (and can't nominate anyone else)

We can draw the model as a flowchart, showing how people move between the different groups,

Here, we've divided nominees into those who accept the challenge (S->I) and those who don't accept (S->R). So, now we can describe the model as a system of differential equations,


Where the factors (α, β, γ) describe what proportions of each group move to where.

The most interesting part is the 'S*I' terms in the first two equations - this basically says that the number of newly infected people is proportional to the number of susceptible people AND the number of infectious people.

This makes the system self-limiting, meaning that the number of infected people can't just grow to infinity - since that's not what we observe. Instead, what we see is an increase to some maximum, followed by a steady drop-off. For example,

We don't have hard data on how many people did the challenge over time, but we can get an idea of what happened by looking at the YouTube search numbers for the phrase 'ice bucket challenge' (above, via Google Trends).

I mean, it seems reasonable to assume some loose correlation between the number of people doing the challenge, the number of videos of people doing the challenge, and the number of searches for those videos. Searches peaked on August 21st, in case you were wondering.


A More Discrete Model

The downside to this 'S*I' term is that it makes the equations non-linear, meaning they can't be solved analytically - that is, you can't come up with an 'exact' equation for the number of people doing the challenge on a given day, for example.

But we can do a numerical simulation, instead.

In fact, this is a more reasonable way of looking at the system, since we're interested in a discrete time step of one day - i.e. the 24 hours nominees have to complete the challenge. For the simulation, we're also going to rounded the numbers of people in each group (after each time step) to whole numbers, since you can't (or shouldn't) divide a person into fractions.

Now, we need to define some parameters. First of all we need to define our start populations. We'll call the initial susceptible population S0 - this could be, for example, the Earth's population (which was around 7.16bn when I ran my simulations). We'll assume that one person is infected as a starting point  - the originator of the challenge (I0 = 1). And we'll assume that no-one else has done the challenge at the start, so R0 = 0.

For the constants (α,β,γ), the easiest one to define is γ - the 'recovery' rate. We assume that after the 24 hour challenge period nominees are no longer 'infectious', therefore γ = 1.

For α and β, we have the total rate of infection/nomination defined as (α+β). We're given that each challenge completer gets to nominate three new people, so we can define (α+β) such that in the first step (when there's only one infectious person) we have (α+β)*S0*I0 = (α+β)*S0 = 3. Therefore (α+β) = 3/S0.

Now, α and β are related to the proportions of nominees that accept and decline the challenge, respectively. So we can redefine the constants as α = a/S0 and β = b/S0, such that (a+b) = 3, or alternatively α = a/S0 and β = (3-a)/S0. So now we can look at 'a' as the average number of nominees who accept the challenge.

So to tie it all together, we have the system of (difference) equations,

And it's pretty straightforward to write some code that'll run through those equations.


So we've got our model, what now?

Having a model is all well and good, but why bother? Well, now we can start asking questions. For example, how fast does the campaign spread? How long will it take for the challenge to die out, and how many people will have taken part by that point? And what happens when we change the number of people who accept the challenge?

First of all, if we run the simulation (with a = 2.5) and plot I(n) - the number of people doing the challenge on any given day - we get something like this


[where I(n) has been normalised so that the maximum is 100, as Google Trends does].

This is about what we expected - a rise and a fall. Though it's worth noting the shape is a little different from the YouTube searches above; the simulation drops off quickly, while the search numbers have a longer tail. This could be because, even after the challenges are done, there's a latent interest in (re)watching the videos. Or it could just be that this model is not entirely accurate...


So what happens when we vary 'a'?

We can start by assuming everyone accepts the challenge (a = 3). In this case, eventually everyone in the world does the challenge - specifically within 21 days of the first challenge. But that isn't very realistic.

So let's say that on average 1.5 or 2 of those challenged accept. In these cases, the ice bucket phenomenon ends before everyone can be challenged. For a = 2, the challenge ends after 50 day, with ~13% of the population going unchallenged. On the other hand, for a = 1.5 the challenge takes significantly longer to end (over a year).

In fact it turns out that for any a < 1.6030165.. the challenge will take a significant amount of time to end - the number of 'infectious' people will eventually reach 1, and stay there until S <= S0/(2a) (remember we're rounding each group to the nearest whole number). For a <= 0.5 the challenge ends almost immediately.

Now, we know that for a = 3, the population S goes to zero (everyone is challenged), whereas for a = 2 the challenge stops before S can go to zero. So we have the question - what is the smallest value of 'a' for which S goes to zero? If you do a bit of interpolating, you can figure out that this critical value comes out at around 2.7405349.. At this value of 'a', the challenge ends after 23 days. In other words, if on average 2.74.. (~91%) of the people nominated accept the challenge, then after 23 days everyone in the world will have been challenged.

From what I've seen myself, it seems like nearer 1 in 3 people accept the challenge on average. So, as viral as the campaign was, it was never going to take over the world.

If you play around a bit, you'll find that the critical values of 'a' are dependent on the initial population (S0). But, as far as I can tell, it's not possible to derive these critical values analytically. (Answers in the comments if you can prove otherwise).


Why wait to be nominated..?

At this point, you're probably thinking this isn't a very realistic model. And you'd be right. So lets make it a bit more complicated.

In particular, lets add spontaneous participation - that is, people who aren't directly nominated, but who see all the other people doing the challenge, and decide they want to take part too.

We'll assume that this participation is proportional to those who have already done the challenge (I and R). So to start with, we need to separate the 'recovered' group into those who actually did the challenge (R), and those who declined (D).

The new model looks something like this,



With difference equations,

In the flowchart, we've introduced this new factor 'δ', the 'inspiration' rate. If we re-define it, like we did for the infection rate, as d/S0, then 'd' can be loosely interpreted as the average number of people inspired to take part by each person who's already done the challenge.

Now we can look at how this 'd' factor affects the things we looked at before - how long it takes for the challenge to end, etc.

Let's start by assuming that the average number of nominees accepting the challenge, 'a', is 1 (out of 3). What value of 'd' do we need for S to go to zero? Do a little interpolating again and you get d = 0.8668653 - that is, if each person who pours a bucket of water over their head inspires (on average) 0.867.. people to do the same, then everyone in the world will participate, within 26 days of the challenge starting. For a = 2, we need d = 0.4046021 for S goes to 0. And so on...

What's a plausible inspiration rate? For a normal person, probably zero, while for a celebrity... I don't know. But on average 'd' is probably very close to zero. I mean, we know for a fact that significantly less than the entire population of Earth has been nominated/taken part in the challenge.

If you plot I(n) for a = 1 and d = 0.1 you get something like this,


In this case, you have that same rise and fall - but this time, you have a longer tail, like we see in the YouTube searches. Is this proof that this iteration of the model is more accurate? Maybe. But as I pointed out before, the YouTube searches don't necessarily accurately represent the number of people taking the challenge over time.


Social Pressure

When a friend does a thing for charity, then publicly calls you out to take part, there's a certain amount of social pressure to comply. I mean, if a friend dumped water over their head (just because), then asked you to do the same thing, probably you'd look at them like they were a crazy person. Anyway, that kind of social pressure is implicitly included in the 'infection' rate - more social pressure, bigger 'a'.

Instead, the sort of social pressure I'm talking about here is the kind that goes: "I should accept the challenge because so many other people have already done it". Or alternatively, "it's okay for me to accept the challenge, since so many other people have already done it".

In other words, the more people accept and complete the challenge, the more likely a nominated person is to accept too. Mathematically, we can introduce this with a term in 'S*I*R'. Or alternatively, we can keep the term as 'a*S*I', but make the factor 'a' a function of R -> a(R).

Anyway, if you're interested, you can try investigating that yourself. Or try adapting the model in some other ways. But beyond a point you can end up complicating a model more than improving it.


A Network Theory Approach to Nominating

So I was eventually nominated for the challenge. But I'm a wimp, so I declined to dump ice water over my head, instead making a donation to the Motor Neuron Disease Association (the UK equivalent of ALSA).

For my nominations, I wanted to try and maximise spread. So I nominated the 3 of the people in my Twitter network who are the most active and well connected, and who I thought would be up for accepting the challenge. Plus, as a secondary effect, I figured they might nominate other people in my Twitter network - the network theory equivalent of wishing for more wishes.

In the end, one ignored the nomination, one acknowledged but didn't accept, and one accepted (in the form of a donation) but didn't nominate anyone else. So I guess that theory didn't quite pan out. But I did at least encourage more charitable giving.


So Yeah

If we had real world data we could maybe test the accuracy of these models. But even without, we can get a sense of how the challenge behaves - for example, we find that there's a critical 'challenge acceptance ratio' that determines whether the viral campaign will go 'pandemic'.

In theory, you could apply this sort of model to any viral campaign, or just anything that spreads 'virally'. The nice thing about the Ice Bucket Challenge in particular, though, is that it has well defined rules for how the challenge spreads from one person to the next.

So, yeah..


Oatzy.


[I need an editor, my pronouns are all over the place.]

Saturday, June 25, 2011

Tumbling, Part One: Some Background

I've talked about Tumblr before, and I've talked about using epidemiology models as an analogue for the spread of 'information' on social networks.

And having spent a fair amount of time on Tumblr recently, I have more insights, and more to say on the subject.


Resummarising

First of all, Tumblr is like Twitter in that a post can be spread from person to person by reblogging (similar to retweeting on Twitter). But more importantly, for a normal person like me, you will find that you'll get far more reblogs than you can ever get retweets. And this is important because it gives you more data to work from.

The key here is that you're typically sharing pictures. And in particular, often pictures that fall into certain fandoms - Tumblr is very good at fandoms. So if you post a picture relating to Doctor Who, then it will attract the attention of Doctor Who fans, who may like it and want to reblog it.


An analogy

Imagine there is some air born contagious disease, BUT it only affects men. Women can't even act as carriers.

Now imagine some guy has the disease. In a normal population, you have an approximately even mix of men and woman, so if the disease spreads, it will spread relatively slowly.

But this guy was just on holiday from his all boys boarding school. So when he goes back after the summer - can you see where this is going?

The fact that it's all boys, who are generally in close proximity, means that the disease is likely to spread through the boarding school like wildfire, infecting (almost) every boy.

(There could still be boys that are resistant to the disease, after all.)

Oh, and for completeness, if this infective guy were instead amongst a group of all women, then the disease wouldn't spread at all. But that should go without saying.


Birds of a Feather

If you are a massive Doctor Who fan, and if a huge chunk of what you post on Tumblr is Doctor Who related, then you are going to attract followers that are Doctor Who fans. You're likely to follow other Doctor Who fans yourself.

So it is, that subscribers to a given fandom with tend to cluster together, forming (relatively) tightly knit communities. If you then post a particularly popular ('contagious') image, then it will spread through the fandom like wildfire, as per the analogy.

Obviously, varying degrees of contagiousness still apply; fandom or not, a crappy picture isn't going to get much attention.

And on the other side of the coin, if you were then to post an image relating to a different fandom, which doesn't strongly overlap with Doctor Who, then the post will spread slowly, if at all.


Big Milk Thing

So I posted this image shortly after A Good Man Goes to War aired. When I last checked, it had achieved a respectable 1,431 notes -> 591 likes, and 840 reblogs.

One of the problems with Tumblr is that it doesn't time stamp reblogs. So, unless you're willing to put in the effort to keep track by hand, you can't get good 'spread over time' data.

So what I did instead was work with 'spread over eccentricity'. Here's a network graph to help illustrate
The nodes are coloured by their distance (eccentricity) from me, going from red (closest), to blue (furthest away) - I'm the reddest dot, in the middle of the circular cluster near the top.

So what we can do is count how many pople are at each distance from me - equivalent to assuming that eccentricity is time linked. Which it isn't, but it's the best we have.

So here's the graph of that spread
And reassuringly, it bears some similarity to epidemiology graphs - it hits at tipping point at the third generation causing a massive burst of reblogs, then slowly dies out; albeit with a secondary hump after generation 5.

You can actually see why the tipping points happen by the clusters in the network graph - the biggest super-node, matt-smith-, exposed the image to a massive audience of Doctor Who fans.

Similarly in generation 5 with fuckyeahdrwho - the second largest circular cluster - though the effect was smaller, arguably because a large number of people had already seen it at that point.


Bar Graph of My Favorite Pies

For contrast, here's what the network graph looks like for a How I Met Your Mother image I posted
In this case most of the spread is directly from me - as the original poster I act as a de facto super-node, as is the case with most posts. The post also doesn't spread very far this time, and in fact starts to die out straight away - as seen in the graph below.
One could argue that, if the image were found by a HIMYM fan blog, it would have seen a resurgence similar to that for the Dr Who image. But on the other hand, I've had Scrubs images reblogged by fan blogs, and still not seen them spread as much as the Dr Who,
In this case, the image was reblogged by fyscrubs - the bottom 'super-node' - and it still didn't spread very far.

So I think it's fair to say that HIMYM and Scrubs fans are (in general) less 'intense', and less numerous than Doctor Who fans. Or at least so on Tumblr.


[edit] - Shortly after I posted this blog, an image I posted last week - that had garnered very little attention until now - hit a super-node. And as with the Dr Who example, it exploded with likes and reblogs. So that's arguably points against a 'spread over time' approach. If only because it's too unpredictable.


So that's that. In Part 2 and Part 3 of this series, I'll be looking at creating a mathematical model of the information spread discussed here.


Oatzy


[For the record, captioned screencaps aren't the only things I post on Tumblr.]

Tuesday, September 21, 2010

The Twitter Virus

Background

So here's what's happened - Twitter has been 'hacked'. Now personally, I don't like such ambiguous use of the word hacked, since it tends to imply that some sort of infiltration or breaking in has taken place. But that's just me.

No. What actually happened is that someone discovered that URLs can be posted that include JavaScript (and in particular the "onMouseOver" function), which was executed when you hovered over said link. Then the Script Kiddies got their hands on it, and all hell broke loose

This is an example of 'code injection' or 'cross-site scripting'- that is, code can be posted to a website - be it by a comment, a status update or whatever - and the site will execute it as if it were part of the site's own code.

For most sites, it's not possible, because comments, etc. are 'sanitised' so that code like this is removed or just displayed as plain-text, so it can't be executed. For example, on Den of Geek they strip away all HTML tags from comments, which has the drawback of disallowing formatting, but gets the job done. But these things can slip through the net from time to time.

One other recent example was on YouTube. YouTube would normally validate it's code by stripping away the "<script>" tag which should have prevented the problem. Except, someone realised that if you start the comment with "<script><script>", only the first tag is stripped away, so the code is still executed.

And this was used to cause all sorts of havoc - redirecting people to porn sites, or adding banners and pop-ups to videos (mostly on Justin Bieber videos), and so on.

In YouTube's case it took about an hour to spot the problem and two more to fix it (apparently). And apparently Twitter has been fixed now (approx. 2hours later). So kudos to them.


The Code

What's interesting about this is how involuntary activating it can be. All you have to do is hover over the link to execute it. Which, to be fair, is both simple and elegant (regardless of how inelegant the code itself looks).

The general form looks like this:

[some URL]/@"onmouseover="[some javascript]

In it's simplest form, the code might be something like this (via Sophos.com):



Which only uses the 'alert' function to create an annoying pop-up with some random message.

As for code that will redirect you to some other site, I can't find an example, but one way of doing it might take the general form but add something like:
window.location.href = [redirect URL]
Which is all very straight forward, very annoying and fairly boring. There's probably other havoc you can reek that will temporarily redesign a person's homepage, graffiti it, or whatever.


The Virus

Now these are personal favourites - the self-retweeting tweets. Literally a self-replicating Twitter status virus.

Here's some example code:



They both basically do the same thing, though the latter does it more elegantly.

The first one finds the 'first text area' - i.e. the status input box - fills it with it's own URL - this.innerHTML - hits the update button for you - ('.status-update-form').submit() - and then darkens the screen with modal-overlay which means that the page itself is unreachable (without reloading) and clicking anywhere reactivate the exploit.



 [Yeah, I wanted to see what it did :p]


For the second one, it gets the element with the tag Id="status". Here's a section of Twitter source code:



So that would be the status input box as with the first example. It's just a nicer way of going about it, in my opinion. But anyway.

And in this case, rather than just straight copying itself, it actually RTs the named user's last tweet (usually the code itself), which again is a nicer way of doing it. And it gives the original poster the credit they deserve (albeit with the potential risk of having their account suspended).

Video of the exploit in action here.

Another variant blacks out the actual status, like so:



And what that means is (a) you don't know what's going on under there, and (b) your curiosity is more likely to get the better of you.

The code is the same as the first of the above, except instead of the the modal-overlay you have "style="color:#000;background:#000; - basically, 'make the status black text on black background'.


One Last Thing

Obviously, the exploit wasn't limited to "onMouseOver". Any JavaScript could've been used (so long as it was 140 characters or less). But nonetheless, the code would need a trigger, and mouseover was one of the best ways of doing it. Others apparently managed to make it activate by moving the cursor anywhere on screen. So yeah.

Personally I'd be interested to see how far this spread. And indeed, given the 'contagiousness' of it, it'd be quite useful for modelling how this - and indeed everything else - spread through Twitter (as I've previously talked about).

Again, a great part of this was the celebs and the connectors; including an early victim, Sarah Brown (Gordon's wife), who has over 1mil followers. And as I said, passing it on isn't as voluntary as RTing.

Why should we have been worried? Imagine the blackout version of the above, but with an added redirect to some malicious site. Yeah. Not that that should be a problem anymore.

I was going to say, if you want to have a play, learn a little JavaScript and have a go. But I guess you can't now. Shame.

Wonder if anyone's tried XSS on Facebook yet...


[Update]

The exploit has definitely been nullified. Twitter explain what happened here.

One of the discoverers of the exploit seems to be @Zzap [source], who had no malicious intent for it. More curiosity. And he, himself, was inspired by the now suspended "RainbowTwtr" - screenshot of how they used it here.

This guy from Kaspersky gives a brief analysis, including a graph of the exploit's growth over time (reaching 93 tweets per second at it's peak). ThreatPost also has some analysis.

The exploitative tweets themselves don't seem to have been deleted by Twitter, but mouseover now does nothing (other than what a link is mean to do).

Anyway. As you were..


Oatzy.


[I'll be honest, I'm not an expert of any sort on JavaScript. But I know enough to get by.]