Wednesday, February 20, 2013

Female Candidates (Part 1)

There has recently been a lot of discussion recently regarding females being under-represented in SU elections on campus.

Quite a lot. All over the place. Honestly. Well, ok, mostly at the Wanderer and all over Facebook.

Frankly, yes, this is an issue. While there are no established barriers to female candidates, we cannot deny that there simply aren't any female candidates this year, and I would further venture that we can't deny that this is a bad thing - this really ought to be addressed.

Unfortunately, everyone seems to have wildly differing opinions as to why this phenomenon is occurring. While y'all can go debate fancy things like sociology, I'm thinking of (over the next little while) taking a look at it from a statistical point of view. Who knows - perhaps this could be similar to the Berkeley Sex Bias case I wrote about earlier, and the problems are more fundamental and basic. Who knows indeed...

In the meantime, here's a fancy thing I learned how to make after looking around on the internet. Enjoy! The size of each block is the size of the faculty, and the colour represents the faculty fraction that is female (with green meaning overwhelmingly female, and red being overwhelmingly male).

EDIT: Source Google Visualization API Sample

Tuesday, February 19, 2013

Voting Behaviour

Hey bloglodytes,

I realize it's been a month since my last post, and I apologize profusely. Not too much, though, since I can check the stats and you have not, by-and-large, been impatiently checking for regular updates. Shame.

On a related note, привіт to the randomly high number of readers from Ukraine!

This is just a short post discussing some of the preliminary results of an analysis I've been interested in doing for a while now. When it comes to SU voting on campus, people often talk about the "Lister vote" or the "Greek vote" as though they're massive and important voting blocks that need to be either wrestled with or courted. (Let me be clear here - when I say "people" talk about it, I really mean about six people, including me, who have nothing better to do and go out for beers altogether too infrequently)

If it turns out that there is any sort of consistency within these voting groups, that would certainly be cool to know as a potential candidate. Unfortunately, the only way to find this out would to look at anonymized versions of full ballots, and compare voting trends within a ballot. For instance, if a significant number of voters clump Lister or Greek candidates together on a ballot, it would be very easy to identify a trend.

Unfortunately the CRO is reluctant to share this sort of information (cough cough), so the closest we can come to this sort of analysis is by attempting to recreate ballots based on the results read-out that's provided following the elections. This is annoying.

With all that being said, the initial (VERY PRELIMINARY) results from last year's election suggest two cool findings about trends within voting.

The first one isn't really that surprising: voters who vote None of the Above as their first choice have virtually random subsequent vote selection. Any given ballot that ranked NotA first and a candidate second was evenly split between all remaining candidates (within 4%). When comparing these subsequent votes to the distribution of first round votes, they aren't even close.

The second one truly only comes from one datapoint, so I'm hesitant to say anything firm on it quite yet. However, an analysis of the Presidential race shows that voters who ranked Farid (Greek) higher than both Colten (Greek) and Adi (not so Greek) were more likely to vote for Adi than voters choosing their first round picks.

Don't get me wrong, in both cases voters choosing only between Colten and Adi preferred Colten by a ratio of about 3:2, however instead of seeing lots of Farid voters subsequently go on to prefer a fellow Greek (The "Greek Hypothesis"), Colten actually got 7% less of the vote share than when looking at first-round votes.

So that's cool.

Friday, January 18, 2013

The Downsides to Hockey Betting

Hey, have you ever "registered or recorded bets or [sold] a pool"? Or have you ever disseminated any information likely to "be of use in gambling, book-making, pool-selling or betting on a horse-race, fight, game or sport"?

If you have, welcome to the dark side of Canada's criminal underworld.

Yum!
Turns out that, under Section 202 of the Criminal Code of Canada, those things are illegal. I'm not an expert by any means, but it kinda sounds like having any sort of unregistered group of people who bet on things, and who discuss information pertinent to their outcome, is illegal. Welcome to two years in jail!

Looking deeper into the Criminal Code, there's another fun point that even if you're betting with a legitimate lottery corporation (like, say, SportSelect), you cannot bet on the outcome of a single game (Section 207 (4) (b)). This was recently debated in the House of Commons where they tried to scrap the rule, but apparently the Senate might be fighting back on this one for a change.

These rules force legitimate lotteries and betting houses to combine bets. SportSelect, for instance, specifies that gamblers have to bet on a combination of between 3 and 6 games, where they will only win if they get all of their predictions correct.

This has a severely negative effect on the gamblers, because even if they are perfectly correct about the probabilities of an event happening, it will happen so much less frequently that they may not be able to afford to keep playing.

For example, let's take a look at the odds that SportSelect is offering for this weekend's hockey games. There are 13 games going on tomorrow, representing the vast majority of teams in the league. Taking a look at these games shows that the average payout for a team to win away is 2.28, win at home is 1.83, and for the games to go to a shoot-out is 5.73. In fact, 11 out of the 13 games were predicted to be won by the home team (but that's not important right now).

Since these are the first games of the season, it's likely that these odds are pretty much more or less pure speculation - lots of these teams haven't even played together yet. That makes it not inappropriate to compare the odds to, say, an average of historical games, and may represent an 'average' set of games.

Which I did.

Last year, which was a regular season, had 1,230 games played over the season. Of these, 24.4% went into overtime, and 60.3% of those went to a shoot out - meaning in total 14.7% of games went to shoot out. Of the rest, an impressive 57.7% were won by the home team (oh ok that previous thing now does look important...). Similarly, the previous year had 12.1% shoot out games, and 53.7% of the rest were won at home.

So let's say on average any given game has a 13.4% chance of going to shoot out, a 48.2% chance of getting won by the home team outright, and a 38.4% chance of going to the away team. If we pretend that those SportSelect odds represent a more or less 'average' set of games, we can compare them like this:

Outcome Odds Payout Expected Result
Home Win 0.4817 1.8277 -12.0%
Away Win 0.3841 2.2769 -12.5%
Shoot Out 0.1341 5.7308 -23.1%

So there you go - if you knew nothing else about a game other than long term trends, you'd expect to lose between 12 and 23% of your money on any given bet. That would put the SportSelect odds somewhere between a casino game and a lottery in terms of expected money lost.

But wait, it gets worse. Watch what happens when you have to bet on at least three games:

OutcomeOddsPayoutExpected Result
3 Home Wins0.11186.1053-31.8%
3 Away Wins0.056711.804-33.1%
3 Shoot Outs0.0024188.21-54.6%

Oh dear.

So basically, the way it's set up now dooms you to failure. What makes it even worse is that, even if you could overcome the 30-50% penalty by knowing more about the system than a historical average analysis, you'd still only win on maybe about one set of bets out of every eight or so. Even if you did have a good advantage, it would likely take a really long time to ever see any real profits off of it.

Oh well. Enjoy the hockey, I guess.

Wednesday, January 2, 2013

When Life Gives you Weather Stats...



...make Statsonade!

So here's my dilemma. I have this 'weblog', and it's really cool when people read it. (Not quite as cool as before, because apparently somebody spam-clicked my ads and now I no longer get money. On the other hand - no ads!) By far the most popular posts are when I talk about weather stuff or SU stuff, and as there are no present SU elections to write about and I've been doing weather posts at the end of each season, I may have nothing good to offer for this week.

Instead, I'm going to make statsonade from the stats that life tossed me. Yum!

In my last post, I presented the weather forecast comparison I had for six weather stations in Edmonton during autumn. It was pretty fun, and only one of the six forecasters outright told me I was wrong.

One of the accuracy measures I used in that analysis was the percentage of time that a forecaster was within three degrees of the true high temperature. An alternative way of presenting these results would be to just outright plot the predicted vs. actual temperature results for each station. Maybe it would look something like this:


There are some wicked fun facts from these graphs. In all of them, the red line is perfection, where what you predict is exactly what you end up with. The data was only presented for autumn, and stations are all quite close to perfect, as well as close to each other - the R-squared values range from 0.900 to 0.926. In general it appears as though most of the stations over-predict the temperature when it gets to the higher range

That's all well and good, but what if we wanted to take this a step further? Is there some combination of  stations that gets you better than any individual station? That would be like a weather model or something.

It turns out that you can actually get a marginally better prediction by using a weighted average of the stations. Consider the following:

T = 0.085TAD + 0.301EC + 0.148GB + 0.155WN + 0.483WC - 0.172CTV

After all that work, our R-squared value is a whopping 0.944. Though this method of aggregating weather forecasts is apparently a minor improvement, it's likely not worth it in terms of predicting the weather.

A fun result of the regression suggests that it would be easier to just take a weighted average of Environment Canada and Weather Channel's predictions, as they make up the majority of the formula. What's really strange, though, is that the CTV predictions get factored in as a negative value. CTV itself has a completely respectable correlation between predicted and real temperatures, but for some reason subtracting a weighted version of their numbers improves the overall prediction (when using them and at least two of any other weather station). Mysterious...

Thursday, December 20, 2012

Fall Weather

Hey there!

I realize that technically tomorrow is the end of fall, but seeing as some crazy people think the world's going to end then, I figured I'd get this done now.

Last season's weather analysis seemed pretty popular, and I decided to continue it into the fall. There are two big changes from last time:

  • I added CTV Edmonton Weather as a sixth weather forecaster. Their system is a little bit different from the other five forecasters in the analysis so far; they only give probability of precipitation numbers up to four days in the future, but do a much longer range of temperature predictions. As a result their total score is only directly comparable to other stations for four out of six of the days predicted, as comparing a number to a rainy cloud icon isn't very fair statistically.
  • I changed the way POP scores are calculated. Previously I used a weird system that was more-or-less based on p-scores, but as soon as a station predicts 0% and it rains, or vice versa, their scores are shot. The new system is based on the Brier Score (a system that other people made up and actually use). In this case, a 0% prediction with rain still gives a score of 0, but it's averaged against other scores. 
Anyway, the winner for the wonderfully balmy season of fall is: The Weather Network. (again!)

Scores for fall (out of 100):
Unfortunately for CTV, their score is a artificially lowered compared to everyone else due to the lack of precipitation forecast on the last two days. It isn't that significant of a penalty, though, as the 5th and 6th day forecasts are weighted the least. If we ignore them and only consider four days, their weighted score would become 65.18, much closer to the others.

It turns out that, with the change in POP scoring system, the numbers for fall are significantly lower than the previous numbers for summer. Take a look at this graph:

Not only do all forecasters do worse during the fall, but they also become less consistent with each other.

Some fun facts!

Best high temperature prediction: Weather Network 1-day prediction: 70.60%
Best low temperature prediction: Weather Network 1-day prediction: 78.38%
Best precipitation prediction: Weather Network 1-day prediction: 76.38

Worst high temperature prediction: Weather Network 6-day prediction: 45.70%
Worst low temperature prediction: Environment Canada 5-day prediction: 49.20%
Worst precipitation prediction: Environment Canada 6-day prediction: 54.28

Some graphs!
Again, CTV scores are only directly compared to the others for four days. What's cool to see is that almost all of the stations consistently lose accuracy the farther into the future they try to predict - which, of course, makes intuitive sense. If you're interested in more of a breakdown of how these scores were developed, you can check out these other graphs.



See ya at the end of winter!

Wednesday, December 12, 2012

Predicting SU Elections

Recently, individual bloggers like Nate Silver and Éric Grenier have gained massive (deserved) notoriety for developing statistical models that prove to be very accurate in predicting the outcomes of major elections. If you haven't heard of them I strongly urge you to check them out!

At the end of the last SU election, I posted a very simple regression analysis of a short list of election statistics describing the executive elections on campus. Since then I've added to my model, and I have a good reason to believe it's been made much more accurate.

Spoiler alert: this post isn't going to have any spoilers. I'm not going to tell you any specific numbers. Sorry, potential candidates!

I considered a significant number of quantifiable parameters. I strictly chose to avoid anything subjective (like debate performance, quality of posters, how chatty they are when we hang out), and was able to break the parameters into three broad categories: popularity, experience, and campaigning.

Falling into these categories were measures like Facebook friends and interactions, number of years served on Students' Council or Faculty Associations, and amount of money spent or fines amassed during campaigning.

The coolest result of the analysis was the different impact of each factor. The lowest-weighted factors were the popularity factors (Facebook friends don't appear translate very easily into votes), and the most important factors actually fell under the experience category. This is actually kind of reassuring, especially as it appears to suggest that the elections may be a tiny bit less of a popularity contest than normally thought!

The current analysis uses the results of 30 candidates running for 12 positions over two years (I skipped 2010/2011 because the lack of contested races really messed things up). While this is by no means a conclusive sample size, the fact that the results are so consistent, even between the two years individually, is really promising. Take a look at this graph:


The graph shows the relationship between the predicted number of first-round votes from the model, and the actual number of first round votes from the election. If the model was perfect, all the points would fall on a perfectly straight diagonal line. As it is, they fall on a pretty great line - out of an ideal coefficient of determination of 1.0, the model yielded a result of 0.945. Also, it correctly predicted the winner of each race, which isn't too shabby. I'm personally pretty happy with that result!

So stay tuned during this year's election, because I'm going to try to use this model to predict some of the results. If that doesn't sound fun, then you need to work on your love of stats...

Thursday, December 6, 2012

Why Deal or no Deal is the Best Game Show Ever

Last year the Supreme Court of Sweden declared that poker was both a game of skill and chance. That's ridiculous - EVERY game falls somewhere on the spectrum of skill and chance (maybe apart from games with no skill like War). Even the best of board games depend on some dice rolling or card shuffling, and at some point superior skill can still be beaten by blind luck.

Game shows are similar. On the one hand are shows like Jeopardy where the relative skill of the contestants almost always gets reflected directly in the score, and in the middle we have games like Wheel of Fortune (where the wheel can kill you) and even Who Wants to be a Millionaire (where randomly assigned dollar values and different questions between contestants don't allow for direct skill comparisons).

At the other end of the spectrum (just before Million Dollar Heads or Tails) is Deal or no Deal. Contestants get to point to cases and randomly eliminate dollar values until they either get bribed off by a computer algorithm or stick with the last dollar value that's left. (I mean it's a valuable contribution to society... did that sound sarcastic or something?) Because the choices contestants make in eliminating suitcases are absolutely random, the only impact that contestants really have is when they're presented with the infamous question Deal or no Deal?

Deal!
The game is wonderful because it's essentially a game theory and economics puzzle all wrapped up in one! In fact, the bare-bones simplicity of the game has made it the subject of several research papers that shed some interesting light on the decision-making processes of the contestants. The game can be very easily divided up into four components:

1. The Host
The host is actually the least important part of the show. Literally he does nothing. Moving on...

2. The Suitcases/Models
The fundamental focus of the game, these attractive prospects drive the action and fascination of the audience for an hour at a time. The suitcases are important too. The game starts off with 26 suitcases which range in value from $0.01 to $1,000,000. Assuming any given contestant picked a suitcase at random and stuck with it until the end, then the average value to be won would be $131,477.50. That sounds pretty good, except that the distribution of prize values is so unevenly distributed that half of contestants would get less than $875.

Round one of the game involves picking and sitting on a suitcase, then arbitrarily eliminated six case values before any real decisions are made. Honestly - these choices can only be random and no strategy has any impact. Because of the number of cases being eliminated, by the time the contestant has to make a real choice the average prize money can vary from a worst-case scenario of $13,420.80 (median $350) to $170,916.30 (median $17,500). That's a massive difference, and it's the range of opportunities that contestants could be faced with before any sort of strategy could take effect.

3. The Dealer
After the contestants set themselves up arbitrarily for success or failure, a mysterious algorithm completely real human being offers to bribe the contestant out of the game. This algorithm shrewd businessman only really takes a small number of things into account: making the game entertaining for the audience, and losing as little money as possible.

Now obviously the game show guarantees that every contestant is going to walk away with some amount of money, and as we saw before if every contestant just held onto their suitcase then they'd each get about $130,000. If the show wanted to produce a given number of hours of material per season, then, they would lose the least amount of money if the contestants each played the game for long time. And what's the easiest way to keep people in? Offer them lousy deals.

In fact, an analysis of bank deal offers has shown exceptional cruelty in offering deals. After the first round, the deal that is offered is usually only about 10% of the average of the remaining suitcases. Nobody in their right mind would take that (then again, most people can't take the average of 20 dollar values in their heads and may not know how much they're being ripped off), and indeed nobody on the US version of the show took that deal. In fact, nobody took a deal at all until the offer was at least 50% of the value of the average remaining prizes, which never occurs until at least 20 of the cases have been opened (round 5 of the game). Shrewd indeed. They did, however, end up at around 95% of the value of remaining suitcases right at the end of the game if the contestants stuck around that long.

The only exception is that sometimes if someone does really really poorly on the initial rounds, the offer tends to be a little bit higher just to make the game less depressing.


Poor lad...
 And finally...

4. The Player
The actual variable in the game, the contestants have the opportunity to just mess things at any point in the time. Before talking about what was observed, here's a riddle. What would you rather choose: a deal for $3,000 or a 50/50 chance between $1,000 and $5,000? What if instead it was a deal for  $7.50 versus a 50/50 chance between $5 and $10, or a deal for $875,000 versus a 50/50 chance between $750,000 and $1,000,000?

In each case the deal was exactly the midpoint value between the 50/50 options, so the expected value for either choice is the same. The only difference between the two choices (from a rationality point of view) is the amount of risk a contestant is willing to tolerate - are they willing to accept a low amount of money for the chance to get a higher amount?

Personally I have a pretty low risk threshold - in fact, I'd quite likely be a terribly boring contestant on the show. What was interesting was that the choices of the contestants on the show depended much less on a rational analysis of the situation, and more on how well they'd done previously in the game.

For instance, candidates who had terrible luck in the game (and could theoretically be left with a  $5 or $10 scenario) were the most risky - very few of them took the deal. This is partially due to the fact that quite frankly a $5 difference isn't really much of anything, but more generally also falls under the category of decision-making known as the break-even effect, where gamblers are more likely to choose options that have the possibility (however remote) to bring them above an arbitrary prize value, even if the expected return is low on the choice as a whole.

Contestants with about average luck (similar to a $1,000 or $5,000 scenario) were much less risky. They tended to sit at comfortable levels of money either way, and settling for a guaranteed money value was much more appealing than risking a couple thousand either way.

What was really interesting was that contestants with the best luck were almost as risky as the players with the worst. Even though the potential swing between choices was hundreds of thousands of dollars, most contestants again rejected the deals offered. This is explained economically by a decision-making process known as the house-money effect, where gamblers are more likely to exhibit risky behavior when they feel like they are playing with money that isn't theirs. Rationally, a contestant in this situation has a guaranteed $750,000, and is facing a choice between a deal of $125,000 or a 50/50 chance between $250,000 and $0, but they tend not to look at it that way.

So there you go! Deal or no Deal really is a fascinating show from a game theory and decision-making theory point of view.

PS A lot of people weren't so happy with the Monte Hall Problem I mentioned in a previous post. Superficially, the last rounds of Deal or no Deal may look a lot like the Monte Hall Problem - a contestant has a case (door) chosen, a third case (door) is eliminated, and then the contestant has the choice of switching their selected case (door) for the one that remains. Unlike the Monte Hall Problem, though, there is no advantage from switching in Deal or no Deal, because the one case that gets eliminated in the intermediate step runs the risk of being the highest-value case. Sorry to disappoint.