Showing posts with label OCD. Show all posts
Showing posts with label OCD. Show all posts

Saturday, March 24, 2012

Some Twitter Infographics

I did some stuff like this before. And I figured, while I was updating my network graphs, why not update some of the other graphics?

And it helps that I worked out how to easily extract data from Twitter (see previous blog). The code is here. Again, rate limits apply.


Who Do I Follow?

This is one of the ones I did before - collect together the bios of the people I follow, then make a word cloud (using Wordle)
Basically, I follow a bunch of geeks and writers. Who like 'things'. So really, same as a year and a half ago.

I would point out though that 6 of the people I follow don't have bios, and about 7 just have lyrics.

Data here.


Who Tweets the Most?

These rates are worked out as (total tweets posted)/(total days online). Obviously, the actually post rate will vary over different time scales..
Bubble chart (made with ManyEyes) - bubbles sized by tweet rate (the numbers on some of the bubbles).

The graph below gives a better idea of relative rates, and 'rankings' (click to embiggen)
The blue line is actual values.

The orange is a logarithmic trend-line. It's a pretty good fit (R2=0.95); and, loosely speaking, it means ~70% of the tweets in my timeline come from ~30% of the people I follow. [cf: Pareto Principle]

You get similar log-shaped graphs when you split up the genders.

Full data here.


Chattiest Gender?

You can read all the explanation, caveats, etc. in the previous posts (here and here). I'm just going to go straight into the data.

I follow 27 men and 21 women (excluding celebrities, etc.). The stats are as follow:
Men:
Average = 6.21 tweets/day
Standard Deviation = 6.47

Women:
Average = 13.79 tweets/day
Standard Deviation = 14.27
For clarity, here's a  boxplot (made in R)
Basically, the women tweet more on average, and their rates are more spread out than for the men. In fact, roughly three quarters of the men tweet less than half of the women. Also, there's one outlier in the female group.

This is similar to what we found last time; although the women's average and spread aren't quite as high (average: 13.79 vs 19.21), and the men's average has increased slightly (6.21 vs 5.29).

If you take the ratio of the averages, the women tweet 2.15 times as much as the men. But maybe I just follow particularly chatty women..

Here's treemap (ManyEyes), which should give you a better idea of the gender balance (boxes sized by tweet rate)
Specifically, the graphic above is 62.5% purple (female).

Data here.


Where in the World Are My Followers?

The site I used last time doesn't seem to exist anymore. So I'm using MapMyFollowers instead. As the name suggests, these are my followers, rather than just the people I follow. Nonetheless..
Mostly in the UK and the US. As you'd probably expect.

I will point out though, some of the locations are a little suspect. Some people haven't made their location available so aren't included, and others seem to be in countries they couldn't possibly be in. But it's the best we can do.

Here's a zoom in on the UK


What Do I Tweet?

Made with Wordle, with data from TweetStats.

Words are sized by how often I tweet them; and by extension, @usernames are sized by how often I tweet those people.

In fact, here are the people I 'mention' the most (TweetStats)
Couldn't get a good source on who @replies me. That was one of the things Twoolr used to do..


When Do I Tweet?

Twoolr used to be awesome for Twitter statistics. But sadly, when they left beta, they started charging. And their free service went to shit. Luckily, I found TweetStats. Weirdly, it doesn't need you to log-in or anything, but somehow it can pull data on (nearly) all your tweets - beyond the 3,200 limit. Strange.

Here's some more graphs
Basically, I tweet most on a Friday and Saturday, and at around 1-2pm.

And I've never tweeted at 5am. But that's probably because I'm always asleep at 5am
Except that one time I got really drunk. (SleepBot)


How Much Do I Tweet?

This is another one I used to go to Twoolr for. And, to be fair, I still could. But that only goes as far back as April '10, and its graphics aren't as clear. Here's TweetStats again
Like I said before, I didn't tweet much in my first year. In fact, I only posted 36 tweets in all of 2009.

Now, the one problem with TweetStats is that 5 month gap in 2010. Why is this significant? Well, I was definitely tweeting during that time. In fact, by my estimates, over those 5 months I posted 5,724 tweets (~37tweets/day). So those 5 months account for 43% of all my tweets.

See, the thing is, in 2010, I was out of university, single, and unemployed. I posted a total 8,823 tweets - 24tweets/day. Since I've been back at university, that number's dropped to 11tweets/day.

That lull in Summer 2011 was when I was spending all my time on Tumblr and watching classic Doctor Who. Incidentally, I haven't posted on Tumblr since the start of September '11. It's terribly addictive, you see. I wouldn't recommend it; unless you're addicted to Doctor Who and Sherlock, and have lots of time on your hands..


So yeah.


Oatzy.


[Self-indulgent statistics, and pretty illustrations.]

Saturday, January 28, 2012

Simple Harmonic Sleep

So I started using SleepBot Tracker to track my sleep. This is what it's looking like so far
Basically, I'm averaging about 8 hours, except on those two days near the start where I had to get up early for exams.

Now, being a physics student, the graph reminded me of that for damped simple harmonic motion (starting from 20/01, ignoring the first 3 points). And being a crazy person, I decided to try and model my sleep as such.

So what's simple harmonic motion?

Simple harmonic motion is a type of periodic motion with a restoring force directly proportional to the the system's displacement from equilibrium.

For example, a pendulum is a simple harmonic oscillator - it has periodic motion, its equilibrium is the lowest point of the swing, and the restoring force is gravity.

Now, one could argue that sleep is SHM-like - we have some typical sleep length (equilibrium), and if we get too little sleep, then we'll tend to sleep more in response, and vice versa (restoring force). But it's a dubious analogy at best.

A damped harmonic oscillator is one with damping, which tends to reduce the amplitude of oscillations. So, like air resistance in the case of the pendulum, which eventually causes it to stop swinging.

I'm not sure what the sleep-based analogy for damping would be.


There's a standard equation for defining a (weakly) damped harmonic oscillator. It looks like this:
Where:- A0 is the initial displacement, the e bit is the decaying term, gamma is the damping coefficient (which determines how quickly the oscillations decay), cos() is the oscillating term, omega is the (damped) frequency of oscillation, t is time (in days), phi is the phase shift, and C is the equilibrium amplitude.

So working out the variables from the data, the model equation for my sleep looks something like this
And the graph looks like this
And to prove I'm not entirely crazy, here are the real values and the model values plotted together
Not a bad fit, right? [error 0.19]

In fact, you might notice the real data is still oscillating a little. But the equation outlined above tends to a constant amplitude of 8.3 (no more oscillations).

To account for this, we could add a baseline oscillation term
And fitting the data again, here's what the graph looks like
Arguably, a slightly better fit. [error 0.16]

So yeah. Basically, I'm procrastinating..


Oatzy.


[Just gotta be careful not to hit resonant sleepquency]

Wednesday, December 28, 2011

Gift Wrapping

In these times of austerity, one has to be economical with one's wrapping paper.

Yeah, I know, this would have been much more useful about a week ago. But I just never got around to it. And I use the word 'useful' loosely; any saving one might get from these 'techniques' will most likely be insignificant.

Nonetheless...


Rectangles

All gifts - no matter shape or size - can be wrapped with a rectangular piece of paper.

This is easily proved. However doing so may not be the neatest or most efficient approach. For example, it's easy to efficiently wrap a cuboid (book, DVD, box) with a rectangular piece of paper, it's less easy to wrap a bike (efficiently) with a single, and in this case, very large piece. However, in the case of a bike, one could use several, smaller pieces of paper for more efficient wrapping.

For a ball, one might be tempted to say that a more unusual shape - maybe a truncated icosahedral shell - would be more efficient. However, cutting out such a shape would result in lots of little slivers of waste, and would take a lot of time and effort. Instead, it's usually easier, and not greatly inefficient to just use a rectangular piece, and some creative folding and scrunching.

So here's the point - for a given gift, we can define a rectangle (or set of rectangles) which most efficiently wraps that gift.

For a cuboid of sides L(ength), W(idth), H(eight), the minimum paper would be: (H+L) by 2(H+W), with a bit extra for overlap.

For a cylinder of dimensions H(eight) and R(adius), the paper would need to be: 2PI*R by (H+2R), and a bit.

[Proof that these cuts are optimal is left as an exercise for the reader.]

For other shapes, this may be more tricky to determine, but for the sake of arguing, we will assume we know all the required rectangle's dimensions.


Bin Packing

So for a given collection of gifts to be wrapped, we have defined a set of rectangles. We now have to determine the most efficient way to cut these out from rolls of wrapping paper - i.e. a large rectangle, ~ 1m x 4m.

This is similar to solving a two-dimensional version of the Bin Packing problem.

In the Bin Packing problem, we have several bins (cuboids) of fixed dimensions, and a collection of smaller cuboidal objects. The problem is to pack the objects into as few bins a possible - i.e. find an efficient way of fitting cuboids into boxes.

The problem is NP-Hard, meaning that finding the optimal solution can take a very long time. However, there are algorithms which are fast and give adequately optimal solutions.

One, in particular, is called the First Fit algorithm, and it works something like this:

1) Sort the rectangles into size order
2) Cut out the largest pieces still in the set
2 i) If there is not enough paper left on the current roll, move on to the next/start a new roll
3) Repeat (in descending size order) until there are no more pieces left to cut.

There are also instructions for how to place the rectangles, for example first fit from bottom and right, to left and top.
This article gives a more precise description of a similar algorithm, though that one is defined for a 'roll' of non-restricted dimensions.


Realistically

That's all well and good mathematically. But in the real world, it's likely more trouble to solve the problem than the saving in paper is worth.

However, the general principle of the algorithm can be loosely applied (if indeed it isn't the approach you already use) to improve on efficiency.

In general, you can easily sight-guess which gift in going to require the largest amount of paper. And from that you can apply the First Fit algorithm.

One change I make is, cut out and wrap the largest one left, first. Then wrap the gift(s) which best fits the off-cut.
This is particularly good for rolls, where you want to unroll and use it a 'slice' at a time (rather than unrolling the whole thing).


Anyway, something to think about next time you happen to be wrapping a large number of gifts.

Even better if you can get some of that Tesco wrapping paper with the grid on the back.


Oatzy.


[No wasted paper from my wrapping this year. Just saying.]

Tuesday, November 08, 2011

Sorting DVDs

So there are plenty of ways to arrange your DVDs, if indeed you chose to arrange them at all. The most popular one would probably be alphabetically. Alternatives may include arranging by release date, by director, actor, genre..

At one point there was some logic to the arrangement of my own DVD. But as I got more DVDs and less space, it mostly devolved into piles, arranged roughly based on the order they were bought in, and which were watched most recently.

If I ever get time, I will re-sort them.


I accept that alphabetical may be considered 'best'. In particular, if you have a significantly large collection, it can be the most efficient approach, and makes locating any given film much easier.

Now, ordering things alphabetically is usually considered quintessentially 'OCD'/fastidious behaviour. Though there is dispute over whether liking your DVDs (and things in general) to be in order really is, or should be, considered 'abnormal'.

My point is, I could do much worse.


First of all, I think, in general, films by the same director should be kept together; particularly when it's a notable director, say, Tarantino, Kubrick, Burton.. So you could group and sort groups alphabetically by director.

I would say that films of similar genre/theme/styles should be kept close as well - noting that films by a given director will generally stay close together when this criterion is applied.

And while you're at it, why not take into consideration: actors, writers, based on works by, producers, scored by.. Basically, there are a lot of potential connections which may be considered.


So I'd like to put forward two approaches:-



1) Salesman Sort

For this approach we take into consideration 'most meaningful connections' between films, and based on that, attempt to determine which films most belong together.

So, for a given set of films, we build a network graph where pairs of films are connected if they share a common director, actors, theme, etc.

Doing this for Tim Burton films, based on actors only, will give something like this
[NB/ There may be connections missing.]

Or like this without the actor nodes
The connections between films are then weighted based on how 'meaningful' they are.

For example, Moon and American Beauty both star Kevin Spacey. Fight Club and Choke, on the other hand, have no actors in common, but were both based on books by Chuck Palahniuk. I would argue that the connection in the latter case is more meaningful - and so should get a higher weighting - than that of the former.


Now, the trick is to find a way of translating this graph into an arrangement of films. It turns out, this is as 'simple' as running a Traveling Salesman algorithm on the graph.

The traveling salesman works like this - find a path around a graph which visits every node exactly once, while minimising the distance traveled.

In our case, we actually want to find a path which (effectively) maximises 'distance', since that will lead to the more significant connections being chosen by preference.

In the above graph, the connections haven't been weighted, so we can pick any path - for example
But ideally, the connections would be weighted first so that the path (and by extension, arrangement) chosen is more meaningful.

Once a path is found, this is turned into an arrangement by simply looking at the order in which the nodes (films) are visited.


The biggest problem with this approach is forming the graph in the first place. One possible solution is to write a program which can pull details from, say, imdb to build connection. Then you'd also need to come up with some system for weighting - which may vary from person to person, depending on their particular preferences.. So it's tricky.

By comparison, the Traveling Salesman part is relatively straight forward, given that it already has well establish (if potentially slow) algorithms for solution.

This approach can also throw up some eccentricities. For example, films by a common director may be split up by a film which has few other connection elsewhere in the graph. This is why I would advise 'fine tuning' by hand.

Similarly, it's possible that the addition of new DVDs to the collection may lead to a major re-shuffle being required. Obviously, how major or minor the re-arrangement depends on how well the film fits in with your pre-existing collection.



2) Hierarchical Grouping and Associative Sorting

This is actually the approach I started out with, before it went to shit.

You might start by grouping by director. Then you might group directors by genre/theme/style. And within the bigger director groups, you might sub-group by actor. And so on..

In particular, the 'hierarchy' is constructed so that some groupings get priority over others. For example, films with the same actor might be split up when those by the same director need to be kept together. (This will depend on how you chose to structure your hierarchy).

Then, within and amongst, DVDs and groups can be arranged according to alphabet, release date, genre, .., even by colour of cover. This is where the 'associative' part comes in. Ultimately, it can come down to how your own demented logic associates and puts things together.

When I was sorting mine, I did end up with some fairly esoteric, and sometimes laboured groupings and connections in places. But you can see some overall logic in it - e.g.

Nightmare Before Christmas / ... / Sweeney Todd / From Hell / ... / V for Vendetta / ... / X-Men..

In this case - Tim Burton, Johnny Depp, Alan Moore, Graphic Novels, and so on. Note that there is overlapping between adjacent groups. They were then sub-sorted so as to be grouped by theme, and so that they formed a natural thematic progression from one group to the next (as much as possible).



Or if you prefer, you could just arrange them alphabetically. Your call..



Oatzy.


[To be honest, I wouldn't recommend either.]

Friday, July 01, 2011

T-Shirt Calendar: Revisited

If you'll recall, in early February I created a T-shirt calendar - that is, a 'calendar' where each square is coloured according to the colour of the T-shirt I wore on that day. Previous post here.

So, with half the year more or less done, here's what the calendar is looking like now,
There's not really much else to say on the subject, other than to note that there's very little pattern to it, and that I wear a lot of green T-shirts.

So ultimately we're left gazing upon it's random patch-work glory, and wondering what its purpose or meaning might be.


Oatzy.


[For what it's worth.]

Friday, February 04, 2011

T-Shirt Calendar

When I was putting together the Life by Numbers blog, I was thinking about other things I could track, without the tracking thing becoming too intrusive.

First couple of things I did was signing up for miso to track TV and film viewing habits, as well as RunKeeper to track my movements on foot (which wouldn't be logged with FourSquare). So far, RunKeeper - or perhaps just my phone's GPS - have been a bit of a disappointment. But that's a whole other story.

The other thing I did was to take a picture a day of myself. I had nothing in mind specifically to do with the pictures, but thought why not?

The obvious use is to do one of those time-lapse, slideshow thing. But those are dime-a-dozen. And besides that, the pictures aren't all head on and centred, so you wouldn't really get the same effect.


Details

Instead, I thought about what was actually in the pictures - what could be gleaned from them.

First things that come to mind are (a) when I shave, and (b) when I get my haircut. And they're things I might come back to once I have more than a months worth of pictures; I don't shave that often.

So the only other thing is what I'm wearing.
The above is a sort of calendar - Sunday 2nd of January to Friday 3rd of February - where each square is coloured to match (as closely as possible) the primary colour of the t-shirt I wore on that day.

The last square's crossed out because I haven't decided yet what t-shirt I'm going to wear tomorrow.
[edit] - I wore a green t-shirt.


There's not really much logic to it, beside that when I pick a t-shirt I don't pick one I can remember wearing within the last week. Yes, I have several green t-shirts.

Otherwise, it mostly just looks like Elmer the Elephant. But it is an interesting way of represent the dataset, even if it serves no useful purpose other than to suggest that green might be my favourite colour.


Oatzy.

Saturday, January 01, 2011

Life by Numbers: End of Year Report

General

Age: 21
DOB: 18/01/89
Height: 6ft 2.5
Weight: 11st 10
BMI: 20.8


Lifestyle & Money


Average night's sleep: 8hrs 34
Typically asleep between 2:30am and 11am
Minimum sleep duration: 5.5hrs
Maximum: 11hrs

Job interviews: 1
Jobs: 0
Job Seeker's Allowance received: £1,660.88

Major expenditures:
Acer Aspire 5542 - £430
HTC Desire - £150 up front (+ £20/month)
Two nights at Hotel 53 (Valentines weekend) - £264

Monthly subscriptions, Jan 2010: £49.99
Current monthly subscriptions: £33.98

Total spend on Amazon.co.uk - £231.77

Overdrawn: 3 times
Maximum: -£17.52

Net change in bank balance*: -£59.24
Net worth as of 31/12/10: £715.87

Dog walks: 1
Visits to the gym: 2
1 new pair of Adidas trainers; Used twice.
1 new pair of Converse, black.


Entertainment

Most listened to artist: Pink Floyd
Most listened to song: Change by Karnivool
Most listened to album: Sound Awake by Karnivool

Books read cover to cover: 4
* Chuck Palahniuk - Survivor
* Chuck Palahniuk - Snuff
* Richard Bach - Illusions
* Neil F. Johnson - Simple Complexity

Books started but not finished: 7
Graphic novels read: ~12

Films seen at the cinema: 3
* The Lovely Bones
* Inception
* Scott Pilgrim vs The World

2010 released films seen: 6

Films rented (LoveFilm): 10

Words written for NaNoWriMo: 28,369
Percent to target: 56.7%

Games bought: 2
* Pokémon SoulSilver
* Professor Layton and the Lost Future


Food and Drink

Average cups of tea per day: 3.35
Percent of all drinks that are tea: 53.5%

Percent of all drinks that are alcoholic: 19.7%
Approximate average units per week: 13*

Favourite alcoholic drink: Whiskey (bourbon)
As percent of all alcoholic drinks: 52%

Second most drank: Wine (21%)

Most frequently eaten animals*:
1 - Cow
2 - Chicken
3 - Pig

Favourite meals:
1 - Ham and Cheese Sandwich (4.25 a week)
2 - Bolognese (1 every 8 days)
3 - Mixed Kebab & Chips (1 every ~9 days)

Favourite Snack: Ice Cream
Average bowls per week: 2.65

Beer festivals: 1
Visits to Cadbury Land: 1


Travel & FourSquare

Current Mayorships: 16
Badges: 15

Percent of days checked-in on: 58.7%
Average number of 'days out' per week: 4
Average check-ins per 'day out' - 3.68

Most frequently visited venue: Costa Coffee (Parkgate)
Check-ins per week: 1.6

Most frequently visited franchise: Costa Coffee
Days out that include visiting a Costa: ~67%

Check-in locations (via 4mapper):
Local (red spot=home)


Long distance train journeys: 7
4 x Birmingham
2 x York
1 x Loughborough
Total cost: £127.25

Vintage train events visited: 2
Vintage train magazines appeared in*: several, unwittingly


Twitter

Total tweets (as of 31/12/10): 9091
Days online: 608
Average life-time tweet rate: 15 tweets/day
Average tweet rate over 2010: 30.5 tweets/day
Most tweets in a single day: 93 (on 08/09/10)

Composition of tweets:
34.3% Replies
5.46% RTs
0.35% #FF

Average tweet distribution over 1 day:
New friends*: 28
Total foll/followers (as of 31/12/10): 57/120

Most talked to friends:
@aaangst (3.69)
@PkmnTrainerJ (3.04)
@SallyBembridge (3.01)
@Shinelikestars_ (2.26)
@Aerliss (1)

Replies to the above 5 made up 27% of all my tweets; 79% of all replies.

Tweets retweeted: ~1 in 26
Equivalent RTs per day: 1.17



Notes

* Net change doesn't take into account cash and savings. Net worth does.
* based on an assumed average of 1.5 units per serving. Recomended maximum intake: 21 units/week.
* based on which animal the main meat constituent (of a meal) came from.
* no, I don't know which publications precisely.
* 'friends' here defined as a mutual follow. Though I would consider most of them friends, to varying degrees.


You can compare these numbers with those from August 31st here.

Further stats on individual websites below. Due to various reasons there are limits on what data I actually have. So it's worth pointing out that following websites only cover the last -% of the year,

Daytum - 80%
FourSquare - 74%
Twitter (Twoolr) - 66%

Though the interesting thing about Twitter is that while the data I've got only covers 40% of my total time on Twitter, it covers 81% of my total tweets!

Secondly, some of the details (datum in particular) is loose estimates - things are measured in 'serving' and imprecisely at that. FourSquare only logs places where I can and have checked-in - people's houses, or places not on FourSquare aren't logged.

There are also things I didn't/couldn't track, which I might consider in future. For example local trains and buses I don't get tickets for because I travel for free. DVDs, I can remember which I've bought this year, and films I can't remember precisely what I've seen. Cash transactions I didn't track. And so on.

The question is ultimately whether I care enough about having the record to bother to put in the extra effort. That remains to be seen. It's a very fine balance between detail and sanity.


Oatzy.