Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Saturday, March 24, 2012

Some Twitter Infographics

I did some stuff like this before. And I figured, while I was updating my network graphs, why not update some of the other graphics?

And it helps that I worked out how to easily extract data from Twitter (see previous blog). The code is here. Again, rate limits apply.


Who Do I Follow?

This is one of the ones I did before - collect together the bios of the people I follow, then make a word cloud (using Wordle)
Basically, I follow a bunch of geeks and writers. Who like 'things'. So really, same as a year and a half ago.

I would point out though that 6 of the people I follow don't have bios, and about 7 just have lyrics.

Data here.


Who Tweets the Most?

These rates are worked out as (total tweets posted)/(total days online). Obviously, the actually post rate will vary over different time scales..
Bubble chart (made with ManyEyes) - bubbles sized by tweet rate (the numbers on some of the bubbles).

The graph below gives a better idea of relative rates, and 'rankings' (click to embiggen)
The blue line is actual values.

The orange is a logarithmic trend-line. It's a pretty good fit (R2=0.95); and, loosely speaking, it means ~70% of the tweets in my timeline come from ~30% of the people I follow. [cf: Pareto Principle]

You get similar log-shaped graphs when you split up the genders.

Full data here.


Chattiest Gender?

You can read all the explanation, caveats, etc. in the previous posts (here and here). I'm just going to go straight into the data.

I follow 27 men and 21 women (excluding celebrities, etc.). The stats are as follow:
Men:
Average = 6.21 tweets/day
Standard Deviation = 6.47

Women:
Average = 13.79 tweets/day
Standard Deviation = 14.27
For clarity, here's a  boxplot (made in R)
Basically, the women tweet more on average, and their rates are more spread out than for the men. In fact, roughly three quarters of the men tweet less than half of the women. Also, there's one outlier in the female group.

This is similar to what we found last time; although the women's average and spread aren't quite as high (average: 13.79 vs 19.21), and the men's average has increased slightly (6.21 vs 5.29).

If you take the ratio of the averages, the women tweet 2.15 times as much as the men. But maybe I just follow particularly chatty women..

Here's treemap (ManyEyes), which should give you a better idea of the gender balance (boxes sized by tweet rate)
Specifically, the graphic above is 62.5% purple (female).

Data here.


Where in the World Are My Followers?

The site I used last time doesn't seem to exist anymore. So I'm using MapMyFollowers instead. As the name suggests, these are my followers, rather than just the people I follow. Nonetheless..
Mostly in the UK and the US. As you'd probably expect.

I will point out though, some of the locations are a little suspect. Some people haven't made their location available so aren't included, and others seem to be in countries they couldn't possibly be in. But it's the best we can do.

Here's a zoom in on the UK


What Do I Tweet?

Made with Wordle, with data from TweetStats.

Words are sized by how often I tweet them; and by extension, @usernames are sized by how often I tweet those people.

In fact, here are the people I 'mention' the most (TweetStats)
Couldn't get a good source on who @replies me. That was one of the things Twoolr used to do..


When Do I Tweet?

Twoolr used to be awesome for Twitter statistics. But sadly, when they left beta, they started charging. And their free service went to shit. Luckily, I found TweetStats. Weirdly, it doesn't need you to log-in or anything, but somehow it can pull data on (nearly) all your tweets - beyond the 3,200 limit. Strange.

Here's some more graphs
Basically, I tweet most on a Friday and Saturday, and at around 1-2pm.

And I've never tweeted at 5am. But that's probably because I'm always asleep at 5am
Except that one time I got really drunk. (SleepBot)


How Much Do I Tweet?

This is another one I used to go to Twoolr for. And, to be fair, I still could. But that only goes as far back as April '10, and its graphics aren't as clear. Here's TweetStats again
Like I said before, I didn't tweet much in my first year. In fact, I only posted 36 tweets in all of 2009.

Now, the one problem with TweetStats is that 5 month gap in 2010. Why is this significant? Well, I was definitely tweeting during that time. In fact, by my estimates, over those 5 months I posted 5,724 tweets (~37tweets/day). So those 5 months account for 43% of all my tweets.

See, the thing is, in 2010, I was out of university, single, and unemployed. I posted a total 8,823 tweets - 24tweets/day. Since I've been back at university, that number's dropped to 11tweets/day.

That lull in Summer 2011 was when I was spending all my time on Tumblr and watching classic Doctor Who. Incidentally, I haven't posted on Tumblr since the start of September '11. It's terribly addictive, you see. I wouldn't recommend it; unless you're addicted to Doctor Who and Sherlock, and have lots of time on your hands..


So yeah.


Oatzy.


[Self-indulgent statistics, and pretty illustrations.]

Monday, March 05, 2012

So What Was the Best Day To Go Shopping?

Alright, let's be done with this.

Just a quick reminder - what I did was collect Foursquare check-in data for various shopping centres around the UK, in the hope that the data might show something interesting.

Previous blog posts on this data collecting - Best Day to Go Shopping, Panic Saturday, Christmas Eve.

Anyway, I've been collecting data for over 3 months now. And that seems like quite enough.

Here's a graph of (normalised) averaged check-ins on each day of the week for 4 periods:
DecAv (blue) is 21st Nov 2011 to 18th Dec 2011
ChrAv (grey) is 19th Dec 2011 to 1st Jan 2012
JanAv (orange) is 5th Jan 2012 to 2nd Feb 2012
FebAv (green) is 6th Feb 2012 to 4th Mar 2012

Aside from the two weeks either side of Christmas (grey) - when people, apparently, did their shopping more midweek - the pattern is basically the same.

For further clarity, here's the average of those three averages (excluding Christmas)
And here is the order of days, from least to most busy ('relative busyness' in brackets):

1) Wednesday (1.00)
2) Monday (1.01)
3) Tuesday (1.03)
4) Thursday (1.11)
5) Sunday (1.15)
6) Friday (1.24)
7) Saturday (1.78)

Note that the differences between Monday, Tuesday, and Wednesday are not statistically significant - they're essentially the same, and are likely to be as busy as each other/not noticeably different.

So, to answer the title question - Monday, Tuesday, and Wednesday are the best days to go shopping. At least, in as much as they're the days shopping centres are likely to be least busy. And, as you'd expect, Saturday is, by far, the worst/most busy.

And the last thing to point out is that these are the averages over 20 shopping centres for a ~3 month period - numbers for specific locations, and at different times (eg holidays) are likely to deviate from the averages.

And, basically, that's that.

If you're interested, you can see the raw check-in data here.


Oatzy.


[That was definitely worth the effort.]

Friday, December 09, 2011

Best Day to Go Shopping

The Question

As the title suggests - which is the best (least busy) day to go shopping on?

Or more generally, how do crowds at shopping centres vary over time? Which days are busiest, or least busy? Are the shops getting busier as we get closer to Christmas? Less busy?!

It an interesting question, and one that's probably been looked into before. But still, I had an idea and I'm running with it.

[Feel free to skip straight to the results if you're not interested in statistics and the likes..]


Data Collecting

Data collection on this is tricky. Especially if you don't have legions of people to go out and actually count people. What I want to do is extract numbers with the minimum of effort.

So here's the game - Foursquare.

If you're unfamiliar, Foursquare is a 'social game', for which you 'check-in' to locations and earn points and badges accordingly. It's also good if you're an obsessive types who likes to keep track of where they've been.

So the idea is this - some subset of shoppers will be Foursquare users, who will check-in when they visit any given shopping centre. If we can extract check-in counts over a given time period, hopefully that can be used as an indicator of a place's 'busyness'. Obviously, this is flawed - but more on that below.

Right. So on the Foursquare page for a given venue, there isn't a total historical record of check-ins over time. Instead, what we have is the total number of check-ins at that location up to the time when you loaded the webpage.

What we do, then, is record that number at some fixed time every day (say, midnight). Then the number of check-ins on a given day is the difference between the total at the end of the day and the total for the end of the previous day. Easy.

In fact, to make life a little easier, I wrote this bit of code [python]. All I have to do is remember to run that every evening, and we have our data.


Sampling

There are two possible sources of sampling errors:

1) Location

If we only track one location, we have a very small sample size. That means we're subject to perturbations - for example, a major event like the Christmas lights being switched on - or just general statistical noise. Also, shopping patterns may vary across the country, or depending on how close to a city centre the centre is located, and so on.

It's actually fairly easy to overcome this. First of all, we have this list of the largest shopping centres in the UK. From that list I picked 20 locations to sample. This data can then be normalised and averaged to look for any general patterns that are (relatively) store independent.

Oh, and I should probably mention, since this data is being collected from places in the UK only, patterns may vary for different countries.


2) Users

Using Foursquare data, we're working on the assumption that as the number of shoppers increases (or decreases), the number of Foursquare check-ins will increase in proportion.

This is not necessarily the case.

First of all, we look up the Foursquare user demographics. There is no one source of definitive data on this (that I could find). But to get a general idea, there is this, based on a survey of BART travelers.

Obviously, this demographic source is for users of an American transport service, but I'm assuming it's representative of Foursquare users in general.

From this we see that the typical user is most likely male, age 25-34. Or to put it another way, women, young people, and old people are under-represented. And from my experience, it seems like women and old people are the most common shoppers on weekdays.

So this may introduce a disparity between the data and reality. But it's not one we can really do anything about (without seeking an alternative source of data). So, as long as there is a general size proportionality between shoppers and check-ins, we'll consider the data acceptable.


Results

By far, Saturday is the worst day to go shopping (in terms of crowds). But you already knew that.

So I have my data for the last 3 weeks, for 20 shopping centres across Britain. Here's the raw data, for if you're into that sort of thing.

I worked out the check-ins for each place on each day, then normalised by shopping centre - so that the total number of check-ins for each shopping centre over the three week period now adds up to 100. I then averaged these 'norms' across all shopping centres.

Here's what those results look like.
In fact, I went back and 'tidied up' the data, removing venues with less than 100 check-ins total during the recording period - since their sample sizes were maybe too small for any patterns to be statistically significant - and removed a couple of anomalies (one place ended up with negative check-ins).

This is what the tidy plot looks like
[Updated since original post]

Pretty similar, but some of the bars are now closer together (removes the anomalies from Sunday and Wednesday).

So from the results above, the order of days, from least to most busy, seems to be:

1) Monday
2) Thursday
3) Tuesday
4) Wednesday
5) Sunday
6) Friday
7) Saturday

But note, it's pretty close amongst the top 3 least busy days.

It shouldn't be too surprising that weekdays are less busy than weekends - what with people working.

As for Sunday being so low compared to Friday and Saturday - well, that might be the result of Sunday opening hours.

For example, Meadowhall has typical opening hours of 9am-8pm, but on Sundays it's 11am-5pm -> 11hrs vs 6hrs. So maybe it would make sense to re-adjust accordingly. But deciding how, exactly, to re-adjust is tricky. So we'll just leave it be.

NB/ Wednesday, week 2 maybe distorted due to public sector strikes - there certainly appeared to be more people on the train. But but a lot of that increase was from children (see sample bias above).


Of course, each shopping centre is unique, and there will be variation as to which days are best and worse for each. As an example, here's what the (non-normilised) plot looks like for Westfield London (the most checked-in to shopping centre by far)
Again, pretty similar to the average. But in this case, Wednesday is less busy than the average, and Thursday more. And Friday has that weird dip in week 2, bringing its average down.


One last thing I'd like to point out - notice there is no particular week-on-week trend. That is, the number of check-ins isn't (on average) increasing as we get closer to Christmas. Or decreasing for that matter. Which is, perhaps, not what you'd expect.

But maybe that will change within the next couple of weeks. And certainly after Christmas, when the January sales kick off. Maybe.


This is an on-going project - bear in mind, this is only 3 weeks worth of data, so it may be too soon to draw any solid conclusions - but I will keep you posted. Maybe I'll do another post just after Christmas, or after New Year's. At any rate, I'll tweet it when I do.


Oatzy.


[There's always online shopping..]

Sunday, September 11, 2011

Picking and Scoring Words for Hangman

Hangman, for those unfamiliar, is a popular way for primary school teachers to teach spelling in the guise of a fun guessing game, and a way for secondary school teachers to pass the time at the end of term when they can't be bothered with teaching.


It Works Like This..

In a dystopian future where the powers that be like to play deadly games with their political prisoners, Hangman sees one such prisoner, 27 year old Adrian Silk, standing on a platform in front of a cheering crowd; his crime - logical thought.

Adrian is presented with an unknown word - 9 letters, title of a 21st century movie - and his task is to determine that word by guessing letters.

With each correct guess he gets closer to discovering the mystery word and winning a stay of execution. With each incorrect guess, another piece of the gallows is built. If he can't reveal the word before the gallows are fully constructed, he'll be dancing a dead man's jig for the crowd.

He starts by guessing the vowels - A, E, I, O, U - two incorrect guesses, two pieces of the gallows built: I__E__IO_

He guesses some common consonants - N, R, S, T - two more incorrect guesses, but still safe: IN_E_TION

Adrian stops to think - "are there any films called INFECTION? There are two, but they're pretty obscure.. Wait..!". And he hazards a guess - INCEPTION ..?

He's right! Adrian jumps for joy, and breathes a sigh of relief. And as the disappointed crowd jeers, Adrian is shot between the eyes by the Games Master - found guilty of the logical thought with which he'd been charged.

Still, if you think that's bad, you should see how they play KerPlunk!


Lets Play a Guessing Game

Okay, imagine you were guessing letters for an unknown word. But in this variation on the game, the hangman will only tell you when you've won, or else when you've run out of letters. That's all the information you get, nothing else.

The simplest approach to finding a word would be to just work your way through the alphabet - 26 letters, a maximum of 26 guesses to uncover any word.

But there will be a limit on how many incorrect guesses you can make - usually around 10. This means that, unless the word you're guessing contains 17+ unique letters this probably isn't the optimal approach.

So with no feedback and a guess limit, all you can do is guess letters 'at random'. But some letters appear more often in the English language than others. So you can make educated guesses, and go for the more common letters first: E, T, A, I, O, N, S, ...

If we assume that the probability, p(ci), of a person guessing a given letter is roughly equal to that letter's frequency in the English language, then the probability, big P, of that person guessing a correct letter in a word equals the sum of the probabilities of each unique letter in that word.

For example, in MISSISSIPPI, the unique letters would be {I,M,P,S}, so

P(MISSISSIPPI) = p(I)+p(M)+p(P)+p(S) = 4.025% + 2.406% + 1.929% + 6.327% = 14.687%


What this immediately suggests for choosing words is:

1) Words with 'uncommon' letters are better

Uncommon letters have lower probability p(ci), so will make totals lower. E.g. C is less common than H, so (in the game described above) CAT is harder to guess than HAT

2) Words with fewer unique letters are better

If there are fewer (and smaller) targets, it's harder to hit one of them by just firing at random.


Parting Words

So what about in a real game, where you get lots of feedback - which letters are right or wrong, where each letter appears and how many times.. This makes deduction a lot easier.

Have you ever tried to cheat at a crossword? There are loads of 'helpers' online - put in the pattern and see what possible words match it. Well for hangman we take that a step further.
Say our pattern is E_UA_I__; there are 5 words it could be - EQUALITY, EQUALING, EQUALISE, EQUATION, EQUATING

First of all, EQUALISE is not a valid solution because it introduces an extra E. And if we'd guessed all the vowels to get that pattern, we can also rule out EQUATION since we've already determined that O doesn't appear in the solution word.

So to put it more concisely - we want to find words that match a given pattern, eliminating those which contain extra letters already guessed.

In the case above we're left with 3 words - EQUALITY, EQUALING, EQUATING - and which of the three the correct word is can be determined by guessing the letter T:

1) E_UA_IT_ which can only be EQUALITY
2) E_UATI__ which can only be EQUATING
3) T doesn't appear in the word, in which case it can only be EQUALING

The fact that T is the most common consonant probably means each of those 3 words are poor choices for hangman. This is an example of a low entropy pattern.

Imagine, instead, that you're presented with a 3 letter word, and you've gone through the vowel and found you're left with the pattern _A_

What could that be? Maybe you guess T next and get _AT - well that could be BAT, CAT, HAT, RAT, .. Or maybe T is wrong, then the word could be MAN, CAN, BAY, RAY, WAX, TAX, .. how many guesses have you got left?

In fact, there are 179 possible words which match this pattern. This would be a high entropy pattern.


So, more tips for picking good words:

1) Pick obscure words

Even if there is only one word a pattern could fit, if the opponent doesn't know that word, then they are forced to keep guessing letter-wise. Which makes things a little more tricky.

2) Pick shorter words

There are more 3 and 4 letter words in the English language than 8 and 9 letter words.

3) Avoid too many repeated letters

For example, MISSISSIPPI only has 4 unique letters, but once you've guessed I and S, you've uncovered 73% of the word, and there's really nothing else _ISSISSI__I could be.

4) Pick words with high entropy patterns


A Question of Uncertainty

In Information Theory, entropy relates to the uncertainty in a piece of information. It's measured in bits, and, for a set of equally likely outcomes, is calculated as the base two logarithm of uncertainty.

So a coin toss has an uncertainty of 2, because there are two possible, equally likely outcome. That gives an entropy of 1 bit. For a given character in a word, if we assume all letters are equally likely, then there is an uncertainty of 26 - so that gives an entropy of log2(26) = 4.7 bits.

But as pointed out above, all letters are not equally likely in the English language. So the entropy is more complicated to work out: H = -SUM[p(ci)*log2(p(ci))]. Anyway, it turns out the entropy of a given unknown letter is 4.18 bits.

In the case of word patterns, we work out entropy as log2 of number of valid words matching the pattern, since all matching words have the same probability of being correct.

For the low entropy example above, there are 3 possible words that match the pattern so that would give an entropy of log2(3) = 1.58 bits. The 3 letter, high entropy example, on the other hand - with it's 179 possible solutions - would have an entropy of 7.48 bits.


So Where Does This Leave Us?

The idea is this - to come up with a scheme for scoring the 'quality' of a word in terms of how hard it is to guess in a game of hangman.

First, we score our word by its letter, as described above, to get the probability of guessing a correct letter. We then take one minus this to find the probability that a letter guess is wrong - the lower the probability of guessing a correct letter, the higher that word's score.

Entropy has a well define method of measurement in Information Theory, as discussed above.

But if we measure the entropy of a word before any letters are guessed, we find that all words of the same length have the same entropy. So instead, it would be more useful to measure a word's entropy after, say, some common letters have been filled in. Since most people start by guessing the vowels, that is what I'm going to go with.

In some cases, there are so many possible pattern matches, that you'd do better to simply make educated guesses at letters until the solution is found. For this, we work out the letter-wise entropy. Again, before any letters are guessed, all words with the same number of unique letters have the same entropy.

So. Each unknown, unique letter in a word can be one of those not already guessed. So we can work out the entropy, H(alpha), of the alphabet, sans the letters already guessed. Then, multiply that by the number of unique letters left to find, n.

So for example, if we guess the vowels and get the pattern _A_, then the entropy is  2*H(consonants) = 5.62 bits


From a Practical Stand Point

To work out the word entropy we can download a dictionary, then write some code which will find and count the words which match our required pattern, excluding those containing already guessed letters.

We then find the 'smallest maximum' entropy - which of the character entropy and the word entropy is smallest. In most cases it'll be the word entropy.

And finally we multiply that by the probability, one minus big P, to assign to each word a score:

(1-P)*min{H(W), H(C)}

Or you can use this code.

In and of themselves, the scores this gives don't have any specific meaning - the exact score for each word will depend on your choice of dictionary and initial guess letters. But so long as you use them consistently, the scores should stand as an indicator of each word's relative 'quality'.

For example, by this system, CAT gets a score of 5.3, and EQUALITY gets a score of 1.4. So CAT is a much better hangman word than EQUALITY.


So What's the Best Word


JAZZ, apparently. In fact, while researching this blog, I came across a few other people's attempts to find the best words for hangman. Pleasingly, they went with different approaches to me.

In fact, JAZZ never occurred to me when I was thinking up random good words. The best I thought of was WAX. Also reassuring is the fact that JAZZ does score highly under my system - 6.07 - and scores only slightly better than WAX, with its score of 5.92

QUIZ is an interesting word. It contains the two least common consonants and the two least common vowels (P=9.89%). But there are only 12 words that match _UI_, so it has low entropy. Its final score is 4.13 - lower than that of CAT, with its more common letter set (P=20.0%).

Similarly GYM; it has P=6.34% and has no vowels, but there are only 29 words it could be, so has a score of 5.48. Or MY, with P=4.38% but only 5 possibilities, score 3.18. So it's a matter of balance between probability and entropy.

Ideally, I'd have run a whole dictionary through my system to look for a list of best words. But I didn't. You can if you really want. Otherwise, I'd recommend one of the blogs linked above for lists.


In Summary

So there you go. You can score some of your own words if you so wish. But for on the fly word assessment, just remember, pick:

1) Words with 'uncommon' letters
2) Words with fewer unique letters
3) Obscure words
4) Shorter words
5) Words with fewer repeated letters
6) Words with high entropy patterns

And if you're trying to guess a 4 letter word with an A in the second position, odds are someone is trying to outfox you with JAZZ.


Oatzy.


[Full disclosure: If I got the information theory stuff wrong, it's because I only know as much as the one chapter of this book, and odd bits on wikipedia.]

Sunday, August 07, 2011

Quick Look: Don't Blink..

"..Don't even blink. Blink, and you're dead!" threatened the frustrated photographer, following a fifth failed photo..

No, this isn't a Doctor Who thing (sorry). No, this is based on another lost article I read a while back:

How many photos do you have to take in order to get at least one where no one's blinking?


First things first, what's the probability of one person blinking when you take a photo of them?

Well for one thing, that's going to vary depending on environment, lighting, etc. And also on the shutter speed of the camera being used to take the picture.. But for simplicity, I'm ignoring all that.

The average person blinks around 10 times per minute, with an average blink lasting 300-400 milliseconds (call it 350ms). So in any given minute, the average person's eyes will be closed for a total of 3.5 seconds.

We'll say the shutter is open for less than the duration of a blink. So the probability of a persons eyes being closed while the shutter is open -> p = 0.058.

[I'll be honest, I'm not not entirely convinced that's right. But I'll go with it anyway.]

So the probability they don't blink -> (1-p) = 0.942


If you're taking a picture of one person, that's only a 5.8% chance of the subject ruining a photo by blinking. So your odds of a good shot are pretty good.

But if we have a much larger group of people - n = 30 - the probability that none of those people blink while a photo is being taken:

p1 = (1-p)^n = 0.165

That's an 83.5% chance that at least one person will blink. Those odds aren't so good.


So if we take S number of photos, what is the probability that at least one of those is 'perfect'?

This goes back to the methods use in the Law Of Truly Large Numbers post - the probability of at least one perfect photo, big P, is one minus the probability that none of the S photos are perfect:

P = 1 - (1-p1)^S

So if we were to take, say, S = 5 photos -> P = 0.594

Bearing in mind that you have to make these 30 people stand around while you take your however many photos, an almost 60% chance of the perfect shot from 5 tries it pretty reasonable.

But let's say you're the panicky sort, and you want to be 90% certain that you have at least one perfect shot..?

Without going into the nitty-gritty, we can find S for P = 0.9, using some logarithms and algebra thus:

S = ln(1-P)/ln(1-p1) = 12.8 shots

And if you can get a group of 30 people to stand still for 13 photos, knowing that there's still a 10% chance you won't get that perfect shot, then more power to you.


When people know they're having their picture taken, they generally try harder not to blink. Especially if you use a count down. So the probability of blinking, and by extension, the number of photos you'd have to take, drops dramatically. The numbers worked out above are probably more applicable for candid shots.

On the other hand, maybe the flashing going off (if you use one) will cause some people to automatically blink. And, admittedly, I've taken pictures of myself that have still managed to get capture me mid-blink. Though that might be down to a delay between click and shutter.


Of course, in this day and age, of digital cameras with instant preview, you can just keep shooting until you get the photo you want. Not like the dark old days of film cameras and photo roulette...


Oatzy.


[The word 'blink' and its variants appear ~17 times in this post]

[As a random aside, working out the number of pokéball you need to throw to catch a given Pokémon is done in much the same.]

Wednesday, August 03, 2011

The Toilet Seat Conundrum

Gentlemen, do you leave to toilet seat up, or courteously put it down after use?

I read a (somewhat tongue-in-cheek) article a while back, I can't remember where, that explained the toilet seat conundrum in terms of game theory.  As best I can remember, it was quite clever. I was recently reminded of it when reading a Cracked article, and thought I'd try to recreate it, and - as is my wont - take it a step further.

Game theory, for those unfamiliar, is an area of maths/economics that studies 'competitive' interactions, "in which an individual's success in making choices depends on the choices of others".


Preamble

The (average) probability of the gentleman needing to 'sit down' when visiting the bathroom, we call p*. The probability of not is (1-p).

If the toilet seat is in the 'wrong position' for a given visit, we call the cost of this c1, and we assume that this cost is the same for both genders. This may not be strictly true.

The simplest cost would be in having to move the seat, typically in the form of mild inconvenience, and the potentially unpleasant experience of having to touch the underside of the seat. I'm also told that there are certain perils in visiting the toilet at the night, if the seat is in the upright position and is required to be otherwise. I can't say this is a cost I've ever experienced.

One might argue that the 'costs' are inconsequential; but for the sake of arguing, they aren't.

For the sheer hell of it, we'll call the woman Alice and the man Bob. Alice and Bob have been in a relationship/living together just long enough to quarrel over such matters. I suspect, for most people, this is a non-issue; but that's not the point of the post.

There is a third possible game, not discussed below, in which Bob can just leave the seat down at all times. In this case, we have c3, the cost of clean up if Bob's aim isn't quite up to scratch.

Oh, and there is a fourth game, where the default position of the the seat is upright. This is the worst possible game for Alice, and is only the best possible game for Bob if p<0.28. I can only imagine this game working in an all male household, and even then (a re-adjusted) Game One works out better.


Game One - Leave It As Is

Probability of the seat being down is the probability of Alice being the last to visit the lav plus the probability Bob was the last and left the seat down. Probability the seat is up is 1 minus the above.

Here's the cost matrix for this game
Where cost is c1 multiplied by the probability that the seat is in the wrong position

To get the total costs to Alice and Bob for this game, we work out

(Probability the seat is down x the cost if seat is down) + 
(probability seat is up x cost if seat is up)
And we can work out the ratio of costs
B:A  =>  2p+1 : 1

In the extreme case, where p=0 (Bob never poops), their costs are equal. But in all other cases, Bob's cost is greater than Alice's.


Game Two - Return to Default

Default meaning the seat is always returned to the downright position after use. The probability of the seat being up is always 0.

Here's what the matrix looks like
In this case, Alice incurs no cost. Bob, on the other hand, incurs double cost - when he needs to urinate, he has to move the toilet seat twice: up before use, and down afterwards.

It's obvious that Alice, once again, fairs better than Bob.

Game Three, mentioned in the preamble, works the same as this, but with 2c1 replaced with c3. Alice still comes out better though. Unless she doesn't like the thought of sitting on a toilet seat that's (potentially) been peed on - even if it is cleaned - in which case, there's some abstract cost to her.

If she doesn't mind, then which of games Two and Three Bob would prefer depends on which is smaller: 2c1 or c3.


Lowest Costs

First of all, we note that in both games Alice comes out better than Bob - incurring a lower cost in both cases. That said, Alice does better in the latter game, incurring no cost at all in that one. So Game Two is preferable to Alice.

But what about Bob?

If we take the cost ratio of game one to game two for Bob, this is what we get
B1:B2  =>  2p+1 : 4

In the extreme case of p=1 (Bob never urinates), Game Two incurs a greater cost for Bob (3:4) - and by extension, Game Two always incurs a greater cost to Bob.

THIS is where and why the conflict arises.

Alice prefers Game Two, Bob prefers Game One.


Tipping the Scales

So Alice would prefer to play Game Two, but she has to encourage Bob towards it. So Alice introduces a new penalty - c2 - for Bob leaving the toilet seat up.

The cost will typically be something along the lines of a bollocking, silent treatment, arguments, or whatever.

So what we do is this - the odds of Bob leaving the toilet seat up, and Alice being the next to use the bathroom -> (1-p)/2

Multiplied by the cost, c2, and added to the pre-existing total cost for game one
Now, we - or rather, Alice - wants the cost to be such that Game Two is preferable, i.e. B'1 > B2

Rearranging and simplifying, we get
c2 > 4c1(2p+1)

However, if Alice were feeling kind, she could introduce a 'reward' for putting the toilet seat down, instead. It works effectively the same - barring psychological, carrot/stick considerations.

For this, Alice would have to offer a reward, R, with
Arguably, Bob could introduce a new cost - or enticement - himself, to 'persuade' Alice towards Game One. But TV leads me to believe that this is seldom thought of, or executed.

This might be because Alice has more to gain/lose - in as much as, Alice can avoid any cost by 'playing' Game Two. Bob, on the other hand, incurs some cost in both games.

You can draw your own conclusions on that one.

In terms of a co-operative solution, if we add together Alice and Bob' costs in each game and compare, we find that Game One has a lower total cost than Game Two.

So one could argue that Game One is better overall. The challenge, though, is convincing Alice that that is the best solution for both of them, given that, from Alice' point of view, she does worse in Game One.


Casino Bathrooms

So this is all well and good, but it's kind of a specialised case - the situation of a house with one male, one female, and one toilet. In our house, for example, we have two males, two females and three toilets. What then?

There are a few other problems with the probability-based approach, as well. For one thing, it uses an average poop-probability for the gentleman. It also assumes both Alice and Bob use the bathroom about the same number of time during a given time period - whereas some people have more robust insides than others.

So for this, we create a Monte Carlo simulation.

[This is what we call excessive commitment to an idea.]

In the simulation, we create a 'person' object, and assign to them a gender, an average number of bathroom visits, and, for males, an average bladder to bowel movement ratio. To capture the day to day variability in number of visits to the bathroom, we use Poisson distributions.

We also create a 'toilets' set, representing however many toilets there are, and their current states -> 1 = toilet seat up, 0 = toilet seat down. Each toilet has an equal chance of being chosen for use by any given person at any given time.

Each person has a counter, which is incremented when the person in question has to move the seat. At the end, these counters are grouped by gender for comparison.


In the Middle of Our Street

So I created a 'house' of two males, two females, and three toilets (variables chosen arbitrarily). Then ran the simulation for 10,000 hypothetical days.

The Game One simulation gives a result of ~ 3.28 seat moves per male per day, and 2.27 seat moves per female per day. That's a male:female seat move ratio of 1.44.

The Game Two simulation gives a result of 9.11 seat moves per male per day - bearing in mind, men have to move the seat twice per standing visit, in this version - and women never have to move the seat.

Code here.

Fun fact: Without additional costs and rewards, Game One is always preferable to men, Game Two to women. Regardless of the balance of men and women in a house.

So now you know!


Of course, in some cultures the conflict never really arises, since it is 'the norm' for men to sit for all visits to the lavatory.


Oatzy.


*[inb4 shouldn't p be the probability of needing to urinate lol]

Tuesday, July 12, 2011

With Enough Tries..?

Probability is tricky. It isn't always intuitive. Coincidences aren't necessarily as rare or as unusual as they might seem.

I can't remember how I got to it, but the other day I came across the wiki article on the Law of Truly Large Numbers. An interesting idea to say the least.

Then a couple of days later I was looking through one of my books for blog ideas, and came across an essay with an example strikingly similar to that in the wiki article (in never gave it a name).

Coincidence?


So what is the Law of Truly Large Numbers?

The Wiki page gives this description:
[The law] states that with a sample size large enough, any outrageous thing is likely to happen.
The example given on the page is a little inelegant, so I'll go with the (abridged) similar example from the book,
Suppose that a really memorable, once in a lifetime coincidence is one which has a one in a million chance of happening today, and that during any particular day there are 100 opportunities... [T]he chance that one of these coincidences will happen to you tomorrow is 1 in 10,000. Still very unlikely...
[But] the chance that every one of the next twenty years will have no one-in-a-million coincidences for you is.. 0.48, or a 48 per cent chance.
According to this extremely rough and ready calculation, there is actually more than a fifty-fifty chance that in the next twenty years you will experience a memorable one-in-a-million coincidence. This also means that for every twenty people you know, there is a greater than 50% [chance] that one of them will have an amazing story to tell during the course of a year.
Now this is an interesting thought.

And it raises an interesting question - If you play the lottery enough times, does winning eventually become significantly more likely? Inevitable?

It's an often quoted 'fact' that you're more likely to be stuck by lightening on your way to buy your ticket, than you are to win. But what does 'the law' have to say on the subject?


Preamble

For this we're assuming a good old fashion, six balls from a pool of 49 lottery.

Probability of winning the jackpot (matching all six balls) with one ticket is 1/13983816 or about 7 in 100million

If you play two lotteries, then your odds of winning are (Odd of winning the first) + (odds of winning the second) + (odds of winning both).

OR, and this is easier to work out,

Let 'Odds of not winning', q = 1-p(winning)

'Odds of winning at least once in two games' = 1 - (odds of winning neither) = 1 - (q*q)

This can be generalised to 'Odds of winning jackpot playing n games', p = 1 - (q^n)


Round One: Will I hit the Jackpot in My Lifetime?

First of all, odds of winning the jackpot by playing every week for a year

p = 1 - [1-p(winning)]^52 = 3.7 in 1million

So not great. How about if you play ever week, starting on your 16th birthday and giving up (dying) on your 86th. Or basically, playing for 70 years. Probability of hitting that jackpot?

About 1 in 4,000 chance. So still not great.

Of course, if you buy 40 tickets a week, then that gives you a 1 in 100 chance of winning the jackpot at some point in your life. But by that point you're spending £2,080 a year on lottery tickets. The average jackpot would have to be more than £14.6 million for the expected return (prize*chance of winning) to make it worth playing.


Round Two: What About Immortality?

So we've got the equation p = 1 - (q^n)

The question is, can we find n - i.e. the number of games you'd have to play - such that the probability of winning (p) is 50:50

The trick is logarithms, and the formula is

n = log(1-p)/log(q)

So for p = 0.5, n = 9,692,842 games, or about 186,400 years.

For a 1 in 4 chance of winning? 77,363 years

1 in 100 hundred chance?! 2,703 years

Alternatively, to have a 50:50 chance of winning in your lifetime (70 years) you'd need to buy 2,663 tickets a week. Yeah.

Basically, even by the Law of Truly Large Numbers, and immortality, you'd be waiting a ridiculously long time and you'd still be lucky to win.


Round Three: I'll Take Anything!

Now wait a minute, I hear you say, I could still win something by matching 5 numbers, or even 3. Okay, that's a fair point.

So you need to match 3 or more numbers to win something. Probability of winning anything in any given game? ~6 in 100,000

So once again, chance of winning something if you play every week for 70 years? 195 in 1,000

Now that's interesting. That's just short of a 1 in 5 chance. But to be worth playing, the average prize value would have to be ~£18,666. Worth it? I'll let you decide*.

And finally, how long would you have to play to have a 50:50 chance of winning something? ~223 years. Or 45 years if you buy 5 tickets a week.

Which is going to be a real kick in the balls if that something turns out to be £5.


Or To Put it Another Way

* Imagine a game you only get to play once. You pay me £3,640 to play, then you pick a number between 1 and 5. I then generate a random number between 1 and 5.

If the number that's generated is the number you chose then you will win some randomly chosen prize between £5 and £5million; you're more likely to win a smaller prize than a larger one, and you can't know in advance what the prize will be.

Want to play?

If you play the lottery, but answered no to the above, you should probably reconsider.


tl;dr As has been said many times before, your odds of winning the lottery jackpot are catastrophically minute. Even if you were to play every week of your life.


Oatzy.

Friday, April 22, 2011

The Perfect Price

Say you made a thing. You put a lot of time and effort into your thing, and you're so proud of it, you want to share it with the world. But how much should you charge for it?

If you price it too high no-one will buy it. If you price it too low, you won't make a profit. And you spent far too much time and money on your thing to not turn a profit.

So you go to a good friend - who just so happens to work in marketing - and ask him to do a little market research for you. This friend is a pretty cool guy, so he goes out on to the streets and shows people your product, and asks them how much they'd be willing to pay for one. But being an expert, he does it in such a way as to get honest and unbiased answers.

After an afternoon of efficient (and pro bono) work, he comes back to you with good sample of 750 responses. After discarding 250 who weren't interested in your product, he analyses the remaining 500 responses.

And by some pleasing miracle, he finds that the prices these people are willing to pay approximates a normal distribution, with average £10 and standard deviation £2.50


So what do you charge? £10?

The people who said they would only pay a price less than £10 won't buy it, because they're cheap-skates, and who needs their business anyway. But on the plus side, half the people surveyed - 250 people - said they'd pay £10 or more. So you would expect to make about £2,500 from the sample group.

Which isn't too bad. But can you do better?

You decide, because you're a bit of a smart-arse, to work out a function for your expected return for if you were to charge £x.

So what you do first is integrate your normal distribution function from x to infinity. This gives you the shaded-area under the curve - the proportion of your sample willing to pay £x,
You then multiply that by the sample size (500) to get the number of people willing to play that amount, and by £x to get how much money you'd make all together. Easy.

Still with me?

Your resulting function looks like this,
(before being multiplied by the sample size)


erfc is the complementary error function, but you needn't worry about what that is exactly, because that's what Wolfram Alpha is for. So, proud of yourself for worked that out (somehow), you plot a graph of this function giving you a graph that looks like this
And right away you spot that there is definitely a peak on that graph, and know that that would be the optimal amount to charge.
So with Wolfram Alpha' help again, it's a piece of cake for you find that that peak is at x=£7.73.

This is your best price. Which is a couple of quid less than the average your sample was willing to pay. But if you were to charge this amount, 409 people from your sample would be willing to buy your thing - and that would make you a respectable ~£3,162 

Good times!

And now you sit back in your chair and laugh, because a little maths just made you an extra 660-odd quid. Which isn't bad going.


As it turns out, if you ask people what they'd be willing to pay (and if their responses approximate a normal distribution) then the price that maximises profits - the one that balances per-unit profit, and expected sales numbers - is ALWAYS less than the average of what people are willing to pay.

And cinemas - whose escalating prices are discouraging movie-goers and leading to declining profits - could perhaps learn something from this. But probably won't.


Spherical Cow in a Vacuum

The world, as you may have noticed, is not an ideal place. Life is never so simple.

The central limit theorem says a normal distribution will often suffice (for a large enough sample population), but it's not necessarily going to be the best fit. Or it might be that the results from your sample don't scale to the general public.

But much worse than that is people. People aren't rational, people don't necessarily know what they want, people don't know what things are worth, and people are surprisingly easy to manipulate - to an extent, you can effectively tell people what they want to pay; as anyone in marketing will proudly tell you, while grinning maniacally and eying up your wallet.


So in that vein, I leave you with these two TED talk - 

Dan Gilbert on our mistaken expectations
Rory Sutherland: Life lessons from an ad man

Watch them.


Oatzy.


[There are lots of other TED Talks on a huge range of subjects. Most worth watching. Some of them are particularly fantastic. Go explore!]

Saturday, February 19, 2011

Quick Look: Uphill Struggle

When the parent's work schedules both coincide with a school day, it's often my displeasure to have to retrieve the small one from school. And because I don't drive, that's a ~30min round trip, which involves a fairly steep hill.

I don't much care for it.

I didn't track how many times that happened last year. But out of shear curiosity, and because I can, I tried modelling the set up to get an estimate.

You could probably work it out with probability alone. But that's too much like hard work.


Assuming...

Firstly, we assume I can randomise the parent's shifts - that is, there's no particular pattern to them. We also assume that their shifts are independent of each other and of school days and school holidays. That's not strictly true. But since there isn't a strong, underlying pattern linking them, we just ignore it.

Same goes for holidays, which in this model are also arranged at random. We can also ignore date, month, etc. because of the above assumptions and because it's just easier that way.

And finally, we ignore sick days, study days and over-time, since they aren't really predictable, and should hopefully average out over repeated trials.

I've also included a 10% chance that, for whatever reason, I won't have to pick the sister up. That number's just a random guess.


Hand-waving explanation

First you create three sets representing a year's worth of shifts for each of my parent and a year's worth of school. You then randomly delete a year's worth of holidays, bank holidays, and training days from the sets.

Finally, you count on how many days, shifts and school days coincide.

You can look at the code here.

Since this is another random use of the Monte Carlo method - you run the model 10,000 times and then draw a histogram of the resulting counts.


Hey Look!

It's a normal distribution
Which isn't that surprising.

The average is around 32 times a year, with about 78% chance that I'll have to pick up the young 'en between 27 and 37 times in any given year.

Oh, and obviously if you increase the 10% probability mentioned above, the average will get smaller (and vice versa)
And that's pretty much it. Any questions?


Oatzy.

Wednesday, February 09, 2011

Follow Up: Rotten Bias

One of the things that stuck in my mind from the last blog's analysis was 2007
The top two films for RT Score and Audience Score were Juno and No Country for Old Men. No Country got the Oscar and the highest RT score, but Juno got the highest Audience Score - this was one of the instances where the hypothesis of the last blog failed.

Now there was something I thought about in passing when I was writing the last blog, but I didn't bother mentioning it. But I thought I'd revisit and look in to it.

So that 'thing' is sample bias - that is, the people voting on films on RT (and creating the audience score) may not be representative of the general population.

I was originally concerned that the users of RT may be predominantly younger people; so then, their opinion and taste would be over-represented in the score. But I dismissed that idea as being a stereotyped view of internet users.

But realistically, it's best not to make assumptions either way if you can get actual details.


Rotten Audience

The are two sites I look at for site demographics. The first is Alexa, which you may have heard of. It has some benefits, but unfortunately for more detailed analysis, they want your money.

The other is Quantcast, which is a lot better, but sometimes their numbers are only estimates (not the case for RT). This is the one I'm using here. Here's what the demographics look like
NB/ this is US only demographics, which make up ~60% of the total visitors to RT. But the results are very similar for other countries.

Percentages to the left of the graphs are the actual distributions of the users. The 'index' numbers to the right show how those distributions compare to the internet as a whole.


Misrepresented Judges


So the 'hypothesis', or at least one the the hypotheses, of the last blog was that the Audience score could in some way be used as a predictor of Oscar winners, in so much as it reflects the 'wide appeal' of the films.

Of course, this only works if the taste of the voting comity for the Best Picture matches well the tastes of the people visiting Rotten Tomatoes and contributing to the Audience Score.

The membership of the Oscar comity is fairly secret, or at least, undisclosed. But we do know it's made up of ~6,000 industry professionals, so it seems safe to assume they're likely predominantly in the 35-60 age group, with maybe an even male/female split. And this could be a source of the weak correlation, since RT users are predominantly in the 18-49 age group - skewed slightly towards younger ages.

But do the different demographics have different tastes in film?

I suppose not necessarily. There are certainly some overlaps between these two groups. But on the other hand, maybe the difference between which films RT users vote highest, and which films gets Oscars is a reflection of these differences in tastes.

As it is, there's no way to manipulate RT Audience scores to better reflect the different demographic make up, and no way to improve on their predictive power.

So just some thoughts.


Oatzy.

Thursday, February 03, 2011

Revisited: Tweets by Gender

So I first looked at this here, with a follow up here. The conclusion was that among the people I follow, the women do tweet more than the men.

So it's now about 4 months since that last post, so I thought I'd have a quick look at how things have changed.

As a quick recap for those too lazy to click the link above, here are the results from last time (now in graph form)
If you want to know about the technical details, you will have to click the link.

And here are the new numbers
It's pretty much the same deal, but it's there for those interested. And just to further clarify whatever point it is I'm trying to make, here's a boxplot comparison of the numbers (made in R)
The three circles over the female's numbers are outliers. And once you factor out the outliers, you find the numbers are actually quite similar. But the middle quartiles (the box parts) are still slightly lower for the males.

The other thing I did, for the shear hell of it, was the rates for the last 4 months; the other graphs are 'lifetime' rates.

You can get the dataset for all these numbers - as well as totals, days online, changes in rates, etc. - here.


What About Everyone Else?

This isn't something I'm going to try and work out myself, because frankly it's not worth the amount of effort it would require.

So I googled it instead, and found this
Apparently, on average women tweet 12% more than men. That number is based on total number of tweets posted though, where mine are based on rate.

But an article I linked in the first post suggested that men and women tweet at about the same rate.

So make of all this what you will.


Oatzy.

[If you see anything that looks like an error, leave me a comment and I'll look into it.]

Saturday, January 29, 2011

Follow Up: More Graphics

I'm sticking these in a new post, rather than updating the last post (and potentially, nobody seeing them).


First of all, as I said yesterday, the 'entrance' squares of the snakes and ladders don't technically exist; in so much as, as soon as you land on them you move to the coresponding exit. So I updated the network graph to reflect this
But this version isn't quite as aesthetically pleasing as the other one. Still, at least no-one can accuse me of not doing it right.

And as you'd probably expect, the number of curves directed into a given square's node in the above graph is loosely related to the probability of visiting that square

And the other thing is, I wasn't happy with the board histogram - the bar chart of how often you land on each square. So I played with it a bit, to try and find the clearest, prettiest, and most easy to interpret representation of the data.

This is what I came up with
The numbers for entrances and exits are stacked, so you can see how often a given exit square is visited in total, and what proportion of those visits are from moving along a snake or ladder.

For example, there's about a 47% chance of getting to the end via the 80-100 ladder.

The arcs show which squares are linked.


Oatzy.

Friday, January 28, 2011

Snakes and Ladders

Or Chutes and Ladder, if you're American. Yes, I am talking about the board game.
The idea came from having read about people analysing Monopoly with simulations, Markov chains and the likes. You can read about it in Math Hysteria by Ian Stewart or find a similar discussion here.

In the book, Professor Stewart also suggests that the reader might like to try the same for Snakes and Ladders. So I did*.


If you're not familiar with Snakes and Ladders and how it works, you had a very deprived childhood. There's an example game board below
[wikipedia]

I don't know if or by how much game boards vary. But for simplicity, I'm assuming this is a representative layout. Variations aren't likely to dramatically affect the overall behaviour.


Simulating the Game

You start with a 'board' with a value for each square. All squares start with a value of 0. You can simulate the dice roll by randomly picking a number between one and six and move accordingly. Whatever square you land on, you add one to that square's value. If you land on a snake or a ladder, you move along it. Repeat until you get to square 100.

At the end, you have a record of how many times each square was visited.

Here's a network graph of all possible moves.
Over-arcs are forward movements (dice rolls and ladders), under-arcs are backwards movements (snakes).

One run of the simulation is equivalent to one possible game. So what you do, is run the simulation, say, 10,000 times, and measure things like, how many times you land on each square, average or minimum number of turns it takes to get to the end, etc.

There is a way of solving this exactly, and mathematically, as discussed for Monopoly in Ian Stewart's book, and indeed as was done in this paper [pdf]. But this way is much easier.

You can get the code here.


Results

Here's a graph of number of times each square is landed on (out of 10,000 simulations). Snakes are green, ladders are red, arcs show which squares are connected.
And you've got this slightly weird, overall, wave shape. And obviously, the snake and ladder landing squares are outliers - most visited ladders, 36-44 and 28-84; most visited snakes, 47-26 and 49-11.

NB/ technically I shouldn't include the 'entrances' of snakes and ladders in the graph, since when they're landed on, you move straight to the 'exits'. But for interests sake, they are included. They aren't included in the moves count as separate moves.

You can also graph the number of moves it takes to get to the end
The minimum moves is 7, and interestingly, you can also derive that answer from the network graph - that is the absolute minimum.

In this simulation, the mode is 21 and the average is 36.2. The maximum, in this case, is effectively around 215. But theoretically, if you were really unlucky you could keep landing on the same snake over and over; in which case, the maximum number of moves tends to infinity.

And, the last thing to look at is how much more likely you are to land on a snake than a ladder. First of all, there are (typically) 10 snakes and 9 ladders, so the simple answer would be 1.11 times more likely.

But when you go down a snake, as mentioned above, there's a chance you can go down the same snake again (or down other snakes or up other ladders). So it turns out, you're actually 1.23 times more likely to land on a snake than a ladder. Which isn't a dramatic difference; but the important thing is, you need no longer wonder.

On average, you'll land on about 4 snakes and 3.2 ladders per game.


End Games

Say you land on square 97 and roll a four; there are three possible ways to end the game:

Past the End - assume that moving past 100 means you've reached the end (this is what I did above).
This is the easiest way to end the game.

Roll Again - stay in the same square and roll again until you roll a 3 (or less).
This is effectively the same as the above, except (on average) it takes more rolls of the dice (but not necessarily more moves) to reach the end.

Loop Back - move three spaces forward (to 100) then one space back. Repeat until you land exactly one square 100.

This last one takes longer to get to the end, and has a greater risk of you landing on the snakes on squares 95 or 98, sending you back to 75 and 78 respectively.

And that's demonstrated in the graph below - both lines are effectively the same until about the last 25 squares, where the 'loop back' diverges towards more visits (as expected).
You also have the moves distribution,
which in this case is broader and pushed to right slightly; the average now 43.3. Note also that 7 is still the minimum number of moves.

Also, in the 'loop back', the odds of landing on a snake goes up again - now 1.41 times more likely than a ladder. So on average, you'll land on about 5.2 snakes and 3.7 ladders per game.


Triple Sixes

Another rule you could chose to include is - if you roll three 6s in a row, you go back to square 'zero'.

According to the simulation, you'll roll a triple six in about 13.4% of games, on average.

But the triple six seems to have no significant effect on what squares you land on, the number of snake and ladders you'll land on, or oddly enough, the distribution of how many moves it takes to get to the end. This is probably because throwing triple sixes is so rare.


So there you go. Bet you didn't think snakes and ladders could be so complicated.


Oatzy.


* truth be told, this is something other people have already done. But I like to try these things for myself.