Thursday, January 16, 2014

Just your friendly neighborhood stats!

The city of Edmonton is very statistician-friendly. As well as having the marvelous open-data catalogue, they also have detailed data sheets on every single one of Edmonton's neighborhoods

Instead of having a big preamble about how much I like Sim City, I figured I'd just jump right into what I ended up doing with Edmonton's stats. Each of the following maps is based on one of the stats collected, and show some pretty cool patterns. (Note: for all maps apart from road conditions, red indicates high values and blue indicates low values)

1. Household Income
 Income Map for Edmonton



Household income ranges from $22,000 to $146,000 average per neighborhood, and the highest income tends to be mostly focused in the southwest of the city. In the southwest it spans across both sides of the river fairly evenly, though the above-average wealth extends farther down into Riverbend and Ellerslie. On the northeast end of the city, though, you can almost trace the river by the precipitous drop in incomes from south to north. Average incomes don't pick up again until about 153rd ave.

2. Property Assessment Value

Property Value Map for Edmonton
There is definitely a correlation between household income and property value, but it isn't necessarily always present. This is particularly noticeable around downtown and just south of downtown, where lots of young professionals are making a good income but are renting. Average neighborhood property values range from $201,000 to $836,000.

3. Road and Sidewalk Conditions

Road Conditions Map for Edmonton
Admittedly, I have no idea what scale is used to measure road and sidewalk conditions. It seems to peak at around 20, and my guess is that higher numbers are good, but that's about as far as I can figure out. (note, here good roads are blue, bad roads are red).

The map for road conditions shows a non-surprising trend where the roads in the outside of the city (the newest roads) are in the best condition, and the roads on the inside are quite a bit worse. Every now and then, though, an inside neighborhood has particularly good road conditions, likely the result of recent work to fix a problematic area. Note: these values are from 2010 data, so if you feel your neighborhood isn't quite as shown in the picture, it's not my fault.

4. Hospitalizations

Hospitalization Rate Map for Edmonton
My best guess for the meaning behind these numbers is that the represent the number of hospitalizations per 1,000 people per year. In Edmonton, they range from 35 to 180.

Unlike the previous maps, the rate of hospitalizations in Edmonton isn't quite so territorial. Most of the city is sitting comfortable around the average, with a notable exception being immediately north-east of downtown (and one bad neighborhood in Mill Woods).

5. People Older than 20 without Grade 9 Education
Missing Grade 9 Education Map for Edmonton
This stat was actually alarming. Likely because I've spent a lot of time on university campuses recently, I've forgotten that some people don't graduate high school. In Edmonton, this value ranges from 4.3% to 41.3%.

This stat correlates very highly with property values and average household income, which makes sense in a way. It's very interesting to see it laid out like this on a map, and I'll leave you to your own conclusions about the ties between wealth and education are.

6. Unemployment
Unemployment Map for Edmonton

2010 Edmonton unemployment was actually fairly impressive. Neighborhood data ranged from 0% to 7.46%, but with an average of 2.69%.

Certainly a lot of the rich areas have tremendously low unemployment, but the rest of the city appears to vary from the average only in certain neighborhoods downtown, with poorer neighborhoods alternating sometimes from high to low unemployment over a distance of only a couple blocks.



I purposefully didn't include statistics that weren't averages or normalized, because though comparing the number of violent crimes between neighborhoods would also have made for a cool map, the differences in size and population would have made straight comparison a bit more difficult. I also excluded rent costs because those were based on 2006 census data, and they were so low compared to today they almost made me cry.

Neighborhood Awards!
Best place to live: Donsdale
Worst place to live: McCauley
The average Edmonton experience: Keheewin
Silliest name: Gariepy (Gary-Epi? Gary-Pee? Silly Gary...)
Least original name: Anything with "West" in it.

Friday, January 10, 2014

Fun with Evolution

Darwin's theory of evolution suggests that random variation can lead to non-random change in species, where individual organisms that are marginally better suited to their environment have a better chance of surviving and passing their winning genes on to the next generation. From this extremely simple idea, repeated countless times, we get the massive diversity and complexity of life.

While it's pretty cool and poetic and stuff, what's particularly fascinating is how it's been adopted into computing science. An entire branch of problem-solving algorithms exist that attempt to recreate evolution, an when reproduced on a large scale can actually be effective at solving incredibly complex problems.

Evolutionary algorithms are particularly well suited to problems where the connections between variables aren't well understood, and computing shortcuts can't be taken, but where programmers know in general what they're looking for. They are also prone to several issues: they're slow, and they tend to latch onto solutions that are pretty good, but not the best (also known as local maximums).

In the simplest terms, computer scientists will create a "gene pool" full of randomly-determined organisms, with the random genes corresponding to variables to be optimized. If the problem is well-defined, each organism can be evaluated and ranked based on how good of a solution they are, and then (similar to real life), the fittest will get together and have baby organisms, with genes from both parents. The process can be repeated as long as desired until a suitable solution is found.

Like real evolution, the non-random selection of the fittest, with random variation in their genes and starting condition, is expected to eventually produce organisms that are well-suited to their environment, but in this case being well-suited means they are an optimized solution to a problem.

In the real world, designers have used evolutionary algorithms to optimize extremely complex designs like car engines or wind turbine blades, but for fun I decided to do a fairly simple test on Excel as a proof of concept to myself. Instead of doing anything useful or fancy, I decided to try to get Excel to draw for me. Specifically, I wanted it to say "HI" to me (hey, I yell at it enough, figured it deserved a chance to yell back).

In order to do this, I politely asked Excel to create 100 random organisms, with each 'organism' defined as a 24-'gene' series of random values. This resulted in 100 random assortments of 6 rectangles. These 'organisms' could look something like this:


And I wanted to give them the goal of overlapping to trace this:


In order to do that, each of my original 100 'organisms' were given a fitness value based on whether they covered the target areas, but lost points based on how much they covered non-target areas. Then they all got a chance to breed like rabbits, but the fittest ones (with the highest scores) had a better chance of breeding than the others.

The actual breeding process is indistinguishable from zookeepers breeding uncooperative endangered animals - I stuck two randomly (but not equitably) chosen organisms in a room and made them watch videos of other bits of data 'doing it' until something popped out. In Excel terms, though, each of the 24 genes were examined in turn, and each child gene had an even chance of coming from either parent. In order to freshen up the gene pool, each gene also had a 5% chance of mutating.

Breeding continued until I got 100 new little babies, at which point I killed off the old stock and started over. In each generation, the fittest have a disproportionate chance of passing on their genes to the next generation, with the idea being that hopefully each generation is stronger than the previous until the goal is met.

These are the results of the first test, after 1,000 generations:


To use technical terminology, this is extremely unimpressive. Two of the rectangles form the I on the right, the lower right leg of the H is there, but a bunch of dead space in the H is taken up on the top, and one rectangle has wandered off completely (there were supposed to be 6...). Don't even get me started on the green rectangle, I think he's shy.

What was encouraging about this, though, was that the average scores for the population did tend to grow each generation, though they ended up plateau-ing after about 500 generations:


This is what I meant before by 'local maximum' solutions. In order for the score to improve, the purple rectangle would have to change by an amount that's more than what I've allowed mutations alone to cover. Also, during the required changes, it's likely that the scores would decrease (as the purple rectangle abandons those upper H legs), which is resisted from an evolutionary point of view. This actually parallels evolution quite well in that once animals are adapted to a situation they stop changing rapidly (even though they may not be perfectly optimized), until external pressures force them to need to adapt again.

The best way to improve solutions from evolutionary algorithms is to increase the sample size (get more genes in there!) and number of generations, but since those both involved forcing my computer to make funny noises, I decided instead to try a brand new set of 100 random organisms. After 1000 generations, I got:


That's... sort of better, actually. At least all six rectangles made it onto the screen, and all three vertical lines are definitely there. Again, though, in order to get this one to perfectly spell the word, the little purple rectangle would have required a tremendously lucky mutation, which was discouraging enough that instead I decided to try one last new group of 100. Here's their final result:


Actually, that's not bad. I have no idea what the little purple guy (why's it always purple?) is doing, but he's pretty much out of the way, and all the major parts of the word "HI" are definitely covered. Not bad for random number on Excel, eh?

In case you really wanted to play around with this spreadsheet, I've put a version of it here. It's already been seeded with random numbers, but if you open it as a macro-enabled spreadsheet, every time you hit "Crtl+q", a new generation will form (Crtl+w for 10 new generations, Crtl+e for 100, but that'll take a bit of time...). Enjoy!

Wednesday, December 18, 2013

Why I Love the Henday

When I was choosing where to get my apartment, my major concerns were cost, decent neighborhood, and access to the LRT for work. That was pretty much it, and as a result I ended up in (what I consider to be) a pretty great location down by the Century Park LRT station.

I soon realized that, while LRT access was great for getting downtown and to sports games, being as far south as I was ended up being fairly inconvenient for getting around to the rest of town, and I was using the Henday ring road a lot more than I had been at my old home, even just to get to other places within Edmonton.

So I decided to take a look at just how efficient the road system is in Edmonton, and how much the Henday played a role in my life. First of all, here's a map showing travel times for someone who lives downtown:


Living smack-dab in the middle of the city definitely has its advantages in terms of minimizing driving times (note: this is assuming no traffic, which is fairly unreasonable for a lot of the time downtown...). Pretty much anywhere between the Whitemud and the Yellowhead is accessible within 15 minutes by car, and, 54% of the city's area is accessible within 20 minutes. Sherwood Park freeway and the Whitemud really open the city out to the east, too.

On the other hand, here is what a similar map looks like for me:


Though it's still a comparable net transit time (53% of the city area is still accessible within 20 minutes), the covered area is very different. This is hardly surprising, of course - sticking someone out at the end of a city ought to increase travel times. What's really cool, though, is that it takes less time to get to the exact opposite side of town than it does to get downtown, even though it's twice as far away. You can even see the effect of the Henday around St. Albert, where a thin band of green colouring hugs the highway.

The real benefit of the Henday is revealed when I plot the same map, but instead avoiding the use of the Henday if at all possible:


Yikes. Pretty much the only easily-traveled areas of the city are anything south of or connected to the Whitemud. Now only 39% of the city can be accessed within 20 minutes, with some areas taking up to 45, and the Cameron Heights neighborhood is pretty much completely lost to me, even though it's fairly close (as the Henday was the closest bridge to it).

Taking the Henday can reduce travel times for me by up to 35%. That's why I love the Henday.

Tuesday, November 12, 2013

Iveson's Friends

So right before the Edmonton election results came out last week, I was extremely excited to run a full statistical analysis on them. I was really eager to see if the effects of signs and flyering could be quantified.

Instead, our new mayor broke a record getting elected, and won handily. It's much harder to do an analysis when he won every single poll (if we ignore hospital and special ballots). If we compare his results to Mandel's results in 2010, we can see the took the support bases of Mandel's very neatly, and then made massive gains in the north and east:


In fact, the only areas where Iveson 2013 really seemed to lose relative to Mandel 2010 were in Karen Leibovici's ward, in the southwest of the city just north of the river. No such effect is really noticed in Kerry Diotte's ward.

Voter turnout was, again, disappointing this year. Here's how it compares to last election:


















So if an analysis on what campaign variables are most important in the election is now difficult (because, let's face it, no factors really contributed to success of other candidates), what can we do?

I decided instead to take a look at how each voting subdivision voted. First of all, I looked at voter turnout and compared that to total Iveson support:


Two things to notice here: first of all, there's a very slight upward trend, which is common for election winners (after all, if the voting subdivisions with lots of voters didn't like you, you wouldn't be likely to win). Also, the fact that this trend is only slight is a good indicator that the election wasn't rigged. There hasn't been much suspicion that the election was rigged (as far as I know), but a similar analysis of Russian election data suggests that certain trends in graphs like this can indicate fishy behavior.

Much more fascinating, though, is looking at the correlation between councillor support and Iveson support in each ward. Each ward is composed of between 15-19 voting subdivisions, and it's interesting to see where support for the mayor and councillor line up, and where they don't. Here are two examples:

In Ward 2, there's a decently strong positive correlation between Esslinger's support and Iveson's, while in Ward 7 there's a similarly strong but negative correlation between Caterina and Iveson.

Now, I know that people often get all up in a fuss whenever they hear about a "correlation", and I'd be hesitant to draw conclusions from this if it weren't for how interesting an analysis of the results from 2010 were. Take a look at this:

In fact, if we put the correlation coefficients into a table we get:

Councillor Correlation
Henderson 0.95
Krushell 0.93
Leibovici 0.86
Iveson 0.74
Batty 0.70
Gibbons 0.65
Anderson 0.47
Sohi 0.40
Sloan* -0.45
Diotte -0.64
Loken* -0.64
Caterina* -0.91
*=Voted against Mandel on arena deal. 
Colours are more or less arbitrary.


This is very interesting, seeing as three of the bottom four councillors were the three who ultimately disagreed with the mayor on the arena deal. Essentially what we're seeing here is that the neighborhoods that really liked Mandel didn't like Caterina, and vice versa. The opposite happened with Henderson - wherever Mandel was popular so was he.

So is it plausible that councillors whose support correlates well with the mayor's are more likely to get along with him? Sure. I'd be wary of using it to predict how councillors will vote on major issues, though, as all it really indicates is how voters reacted to election promises.

Here's the full list for 2013 though! If Esslinger and Henderson tend to work well with Iveson, and Nickel and Caterina don't, just remember that math said it first!

CouncillorCorrelation
Esslinger0.74
Henderson0.71
McKeen0.71
Knack0.60
Sohi0.40
Anderson0.20
Loken0.13
Oshry-0.04
Gibbons-0.13
Nickel-0.26
Walters-0.35
Caterina-0.57

Friday, October 11, 2013

Fake It 'til you Make It

As of October 11th, the cumulative twitter mentions of the Edmonton election were as follows:
  • Don Iveson: 44.2%
  • Kerry Diotte: 32.4%
  • Karen Leibovici: 20.9%
  • Josh Semotiuk: 2.4%
  • Kristine Acielo/Gordon Ward: <0.1%
Mark Blevis, who has been tracking twitter mentions, is performing this analysis to check whether twitter mentions are useful in predicting the outcome of an election. Personally I'm not convinced (especially since the theory failed in the recent Nova Scotia elections), but I love the spirit behind tracking political statistics. The theory that social media engagement correlates to voter engagement certainly isn't without merit, though I'm sure many other factors are also important.

On the other hand, here are the current number of twitter followers for each candidate (as of 11:00 am October 11th):
(*: Emphasis added)

If we ignore Gordon Ward for a second, there's a very strong (R2=0.96) correlation between followers and mentions. This makes a great deal of sense, and all evidence points to a certain proportion of followers of each candidate being engaged in the election discussions on twitter.

But what on earth is going on with Gordon Ward?

He's only had twitter since just after nomination day and already has over 6,300 followers. Virtually nobody mentions him on twitter, though, and the majority of his posts have barely any interaction with his followers (very few retweets or favorites, for instance).

Take a quick look at his followers, and you might discover a pattern - for the vast majority of them, their ratio of following to follower is extremely high, many of them over 25:1. Many of Mr. Ward's followers have fewer than 10 tweets and are not from Canada. Scroll through a couple of his followers and you'll see what I mean. Needless to say, these are not the accounts of engaged Edmonton voters.

Don Iveson wasn't wrong when he called an election a "communications exercise" - a candidate's ideas surely aren't worth anything if nobody ever gets to hear them. Unfortunately, this often results in election candidates resorting to attacks or half-baked publicity stunts to gain the spotlight and have attention drawn to them. Voters often simply don't have the time to pour over every candidate's platform in detail, and often will only pay attention to candidates who they perceive to be popular.

Just like any other communications or marketing problem, the ability to create a 'buzz' around yourself is key to winning an election - only once people think you're worth their attention will they care about your ideas. In book sales, for instance, it's not uncommon for new authors to hire firms to buy up thousands of copies of their books in order to briefly appear on best-sellers lists. Showing up on a best-seller list gives the author added credentials, gets the book noticed by consumers, and will likely drive future sales.

Books and elections candidates are remarkably similar inasmuch as they are judged often by their covers (though dissimilar in that people actually enjoy movies based on books). Much like publishers buying up copies of their own books, candidates can use online services to inflate their social media presence.

If you want to boost your twitter account by 5,000 followers, you could pay $30 and get them within a week. Or maybe 25,000 Youtube views for $100 here. Perhaps 500 Facebook likes for $42? Just like seeing a book on a best-seller list makes people more likely to pay attention to it, having a lot of twitter followers or Facebook fans gives the impression that a candidate is credible and popular, and can hypothetically be valuable in kickstarting a successful campaign. If Gordon Ward had used a service like this (hypothetically, of course), he should probably ask for his money back - being followed by an army of 6,000 zombie twitter accounts doesn't seem to have gained him much momentum so far in this election.

After I pointed out the zombie twitter account horde, it was mentioned to me that similar shenanigans may be occurring on Facebook. Take, for example, these two candidates from Ward 11:


For reference, here are some mayoral candidate charts:





The grey lines in the chart represent the sum of all new likes over the previous week. For instance, from the week of September 2-8, Mike Nickel received 453 new likes. Impressive. However, from the week of September 3-9, he received 0. The implication here is that, on or around September 2nd, Mike Nickel all of a sudden got ~450 new likes on his page, and then didn't get any more until the lead-up to the nomination day. The graph for Mujahid Chak is similar, with approximately 580 new likes occurring right at the end of August. On the other hand, the values for mayoral candidates fluctuate a bit around nomination day, but don't show any of the sudden changes or plateaus of these other two candidates.

Now maybe this was a case of incredible luck for both candidates. Surely there's a chance that they aren't cheating and buying Facebook likes. Perhaps Mr. Nickel had a tremendously successful Facebook hangout and convinced a bunch of people to like him all at once, or maybe Mr. Chak only started his page on August 30 to an incredible amount of fanfare (and subsequently was ignored by the general Facebook community, judging by the "people talking about this" metric...).

I would love to give them the benefit of the doubt. Really I would. Except that Facebook Graph search is a powerful tool, and lets you take a glimpse into the fans of Mike Nickel. It's not available for everyone yet, so here are some screenshots from a search of Mike Nickel likers (Mujahid Chak's results are very similar) :

(Click to expand)

Admittedly, the first few pages of results are pretty clean - lots of Edmonton citizens, fairly legitimate-looking profiles, etc. After about page 4, though, the proportion of people from Edmonton drops significantly. Apparently Mr. Nickel is supported by people from California, Buenos Aires, Uruguay, Turkey, Tunisia, and Vietnam. Broad support base indeed - in fact, of the people who like Mike Nickel and listed their location, 89.1% of them were listed as living outside of Edmonton. Mujahid Chak's supporters with listed locations were even worse, with 92.9% living outside Edmonton.

I of course am open to an innocent explanation for the international social media popularity of some of Edmonton's election candidates, and I can of course sympathize with the desire to be noticed during an election. Though there's nothing illegal or necessarily improper about artificially inflating your social media presence, I personally find the practice to be deplorable.

Note: I did a quick check of most high-profile election candidates for this analysis, but not as in-depth (as they seemed fine). If you happen to notice any others with irregularities, please let me know!

Wednesday, September 25, 2013

Edmonton Election: Donors

I like stats, and I like elections, so I figured I'd write a bit on some of the statistics coming out of the Edmonton election so far (yes, I know we haven't even had official candidates for two whole days yet but bear with me, this will be fun!).

Under Edmonton election bylaws, candidates are required to make their donor lists publicly available following the results of an election. This year, these stats won't be posted until March 1, 2014, but of course candidates are free to do so whenever they wish.

The information from last election is already available, and is fairly interesting in and of itself. Take, for example, the donation results from the two largest 2010 campaigns, Stephen Mandel and David Dorward. Here's a profile of the donations they received:



For this graph, as you move along to the right with increasing donation values added in, you can see the total sum of all contributions go up, all the way to the maximum allowable donation of $5,000. What you end up getting is a fairly smooth profile until around the $3,000-$4,000 range, where all of a sudden people figure if they're in for a penny they may as well be in for five thousand dollars, and you get a MASSIVE spike at the $5,000 donation end.

Up until the maximum donations, Mandel had almost three times as much money as Dorward, but Dorward ended up bringing in the big guns and amassed $85,000 extra in the $5,000 denominations, bringing their final totals much closer together (but of course, in the end Mandel still beat him by quite a large vote margin...).

These graphs are nice and complete because every single donation is accounted for in the declarations by the candidates. Because of the relatively predictable nature of the graphs, the total donations can be easily approximated by breaking the donation amounts into $1,000 chunks, and multiplying them by the average value of each chunk, like so:

Mandel:

Average DonationDonors$ Expected$ Actual
$100-$1,000$550245134,750115,533
$1,000-$2,000$1,5002233,00033,353
$2,000-$3,000$2,5001230,00032,550
$3,000-$4,000$3,500517,50018,200
$4,000-$5,000$5,000*84420,000420,000
Total
368635,250619,636


Error:2.52%

Dorward:

Average DonationDonors$ Expected$ Actual
$100-$1,000$5505932,45036,320
$1,000-$2,000$1,5001218,00021,500
$2,000-$3,000$2,500410,00010,500
$3,000-$4,000$3,500000
$4,000-$5,000$5,000*101505,000505,000
Total
176565,450573,320


Error:1.37%

*Expected donations in the $4,000-$5,000 category are taken to be $5,000, based on the profile shown before.

Unfortunately, donations that are less than $100 aren't broken down by donor, but Mandel and Dorward received $8,548 and $2,020, respectively. This method of breaking down the donations into categories appears to be very accurate at predicting the total amount candidates received in donations, which is handy because of what you're about to read next!

Two of this year's candidates for Mayor, Don Iveson and Karen Leibovici, have within the last week released some information on their donors. Good for them - they certainly didn't have to do it yet, but it's a nice sign that candidates who pledge to be accountable have already gotten the Ball o' Transparency rolling. Their lists aren't quite as broken down as Mandel's or Dorward's from last election, and instead we are given a list of donors and broad categories that they fit into donation-wise. Again, donations of less than $100 aren't listed.

If we take the number of people in each category and try to back-calculate the expected fundraising values (using the previous methodology), we can do a more in-depth comparison between the two candidates, and maybe get a glimpse at the sort of donations they tend to receive. It might look something like this:

Iveson:

Average DonationDonors$ Expected$ Actual
$10-$100$5523012,65013,893
$100-$500$30012036,000??
$500-$2,000$1,2504050,000??
$2,000-$3,500$2,7501952,250??
$3,500-$5,000$5,000*33165,000??
Total
442315,900318,772


Error:0.90%

It looks like the break-into-categories model for Don Iveson's campaign is surprisingly very accurate. The final sum for the <$100 category was the only one that was given, but even the amount for that category was pretty much in line with what you'd expect.

Leibovici's results are a bit different, though:

Leibovici:

Average DonationDonors$ Expected$ Actual
$25-$100$62.5??????
$100-$1,000$5509451,700??
$1,000-$3,000$2,0002346,000??
$3,000-$5,000$5,000*44220,000??
Total
??317,700365,000


Missing:47,300

The results seem mostly fine, I suppose - at first there's not a lot to really compare it to. What's interesting is that $47,300 figure at the end. I suppose it isn't actually missing, per se, and presumably it mostly belongs to the $25-$100 category (Leibovici's website indicated that $25 was the lowest donation they'd received).

What's fun is that, if it is all from the low-donation category, we'd expect a whopping 756 individual people to have donated in that category, if the same model for handling categories that worked so well for Iveson, Dorward, and Mandel is to work here. This is pretty extreme, to say the least. There are three general explanations for what could be causing this discrepency:

  •  Karen Leibovici has found a lot of small-time donors (who apparently haven't donated in this manner to mayoral candidates before). Perhaps this is the first sign of a truly novel strategy?
  • The donation profile for the Leibovici campaign is absolutely wonky, and consists mostly of  $1,000, $3,000, and $5,000 donations. This seems tremendously unlikely.
  • More likely the quoted figure of $365,000 is either not from the same date as the list, or the list as published is incomplete. The seems plausible since the list was published on September 19th but was titled "September 16th", so perhaps a ~$40,000 or so isn't reported on the list, with $365,000 being all donations as of the 19th.
  • Something nefarious is afoot. (Yes, this is usually my first assumption when a model of mine doesn't accurately predict real life...)
Assuming Leibovici's donors are precisely as they've been presented and follow a similar model as other mayoral candidates, we get this profile:


Again, this graph has less data to work with, so it's far less complete than the results from 2010, but still appears to show that Leibovici gets much more support from large donors (at the $5,000 maximum) than Iveson so far. But with the election only having officially begun this week, none of this really matters, I suppose!

Stay tuned for the next post, where I talk about polling. Yay!

Friday, September 20, 2013

Bieber Fever

Newspapers had a hay-day last year following the publication of a paper out of the University of Ottawa that discussed Bieber Fever. Some articles included:
Of all of these, I'm most disappointed in the CBC - I often pay attention to the CBC, and it's discouraging to know that they might be equally as wrong about other things as they are about this.

What's going on here? A professor from the University of Ottawa published a paper where a new model was developed to look at Bieber Fever, and the paper does indeed include the quote "It follows that Bieber Fever is extremely infectious, even more than measles, which is currently one of the most infectious diseases. Bieber Fever may therefore be the most infectious disease of our time." Oh my god, those newspapers must be right! Science has confirmed our worst fears! This must be backed by hard facts and empirical evidence!

Well, no. The paper appears to be a chapter in a book that examines diseases through various mathematical models. Each of the papers in the book (four of which are written by the Bieber Fever author) takes on a different disease and models it, then examines mathematically the effects of different approaches to the disease, like pulse vaccination or changes in infection or relapse rate. I can't comment on the quality of the other chapters in the book, but they seem to be well-developed and certainly based on real diseases. 

The Bieber Fever paper is a little bit different though. First of all, it is clearly written in a tongue-in-cheek manner that I think flew over the heads of most major newspapers. The humor and sarcasm actually make it quite an entertaining read, and if the piece was written as a humorous look into a creative way of adapting a disease model (which is my suspicion), then it could certainly be a fun case study for biology or math students. It is definitely not something worth raising alarms over in newspapers, though, as the model's predictions aren't validated against any actual statistics and its math is misleading, allowing them to draw this ridiculous comparison to measles that grabbed newspaper attention.

Mathematical disease modelling is a pretty cool field. The most basic model that can be developed is an SIR model - a population is divided up into three groups (Susceptible, Infected, and Removed), and people move through the groups depending on disease parameters and the size of the groups at a given time. For instance, if a lot of people are Infected, the chance of a healthy Susceptible person getting infected is quite high (perhaps due to lots of people sneezing on them), but as more people are Removed (happily by recovery and immunity, or sadly by death), it may become harder for the disease to propagate. 



In this model, βIS represents the rate that healthy people become sick - effectively, it is the chance that in a given time a Susceptible person will encounter an Infected person, multiplied by the chance that that encounter will transmit the disease. On the other end, γI represents the rate at which sick people become healthy, effectively the number of Infected people divided by how long it takes them to get healthy (or die, I suppose).

As long as the rate of people becoming sick (βIS) is larger than the rate people are recovering (γI), then the disease will reach an epidemic of some type - otherwise it will quickly die out. For simple models, the ratio of these rates is known as the Basic Reproduction Number (R0) of a disease, and correlates to the number of new diseases a sick person will cause. This is pretty easy to visualize - if the ratio R0 is bigger than 1, then by the time someone recovers from their illness they’ll have spread it to at least one more person, and the disease will grow. If you're unlikely to make someone else sick when you fall ill, the disease’s R0 will be less than 1, and the disease will go away without much of an outbreak. 

For reference, the flu typically has an R0 of 2-3, HIV is around 2-5, Smallpox is 5-7, and Measles is 12-18. For every person who got Measles, the disease was so infectious and you had it for long enough that you were expected to transmit it to between twelve and eighteen people before you either recover or die.

Frightening stuff. Fortunately, analyses of diseases with these mathematical models shows that as long as a certain proportion of a population is immunized by vaccine, epidemics can be avoided. That proportion needed is (1-1/R0) - so a typical flu needs 60% immunization to prevent outbreak, and measles needs over 90%. If you're still unsure about getting a flu shot, just remember that if a population doesn't hit ~60% immunity, it is very much worse off for those who don't have the vaccine or who are otherwise susceptible.

The Bieber paper develops a more complicated mathematical disease model. It looks something like this:


The author, Robert Smith? (not a typo), proposed a model where media effects have a large impact on the disease. Positive media (P in the picture) can increase the rate at which healthy people become Bieber-infected, and can also make recovered individuals susceptible to re-infection, and Negative media can heal the sick or immunize the susceptible (how miraculous).

Using the numbers that Smith? has in his paper, the spread of Bieber Fever in a typical school of 1,500 students would look something like this:



After about 2 months, the system reaches an equilibrium with about 85% of people being Bieber Fanatics. The paper makes a couple of assumptions: first of all, people are assumed to "grow out" of Bieber Fever after a period of two years. People are also expected to interact with everyone else in the population at least once a month, and have a transmission rate of 1/1500. This means that the average infected person will infect 1 person a month for 24 months, giving Bieber Fever an R0 of 24.

SWEET MOTHER OF GOD IT'S WORSE THAN MEASLES!?!?

Not even a little bit! The transmission rate is absolutely just assumed out of nowhere - no stats, evidence, or explanation given. Similarly, the length of the disease is made up, with the explanation "But let’s be honest, we all know which one it really is, don’t we?" (Smith?, p. 7). Essentially, the authors were given a calculation where they had to assume three numbers and multiply them together, and newspapers are surprised that the answer to the multiplication was high. Even the mechanics of the positive and negative media effects are questionable, though the model they developed could help provide insight into other diseases with relapse mechanisms.

The paper is cute, clever, and provides a mathematical analysis of a convoluted set of differential equations - for all of these things it serves a nice purpose as a tongue-in-cheek entry into a textbook examining mathematical modelling of infectious diseases. But newspapers taking essentially the result of an unfounded set of assumptions out of proportion and reporting them as "Science Confirms!" will always annoy me to no end.

One last thing. This is what a graph would look like if the same school was hit with measles:



Now that's an epidemic - three people sick can infect up to 1,200 in less than a week. Remember this when deciding whether or not to immunize your baby.