Monday, September 24, 2012

Summer Weather

Weather forecasting is insane.

As a career I couldn't even imagine how un-rewarding it is - you could pour hours and hours into developing new algorithms that only get tiny increases in accuracy due simply to the massive complexity of the system you're trying to model. When you're right people take you for granted, and when you're wrong you take a lot of blame.

That being said, a while ago I noticed that sometimes different weather forecasters will predict radically different weather for the same day, given the same data. Also, I noticed that on Monday the weather for the weekend could be substantially different than the forecast from Friday. These are all fair differences - tweaks to models could cause differences of opinions between meteorologists, and the closer your prediction is to when you make it, the more accurate we'd hope it would be.

I was curious as to how much of a change there would be, though, which is why I decided to keep track of it. Since the beginning of June I've kept track of the six-day forecasts for High temperature, Low temperature, and Probability of Precipitation for five different forecasting stations: timeanddate.com, Environment Canada, Global Weather, the Weather Network, and the Weather Channel. Environment Canada, Global, and the Weather Network were chosen based on the sites visited most frequently by myself and my friends, the Weather Channel was chosen as it is the basis of Yahoo! weather, and subsequently the commonly-used Apple weather app, and timeanddate.com was chosen because it's a large multinational site. All stations were chosen at the Edmonton downtown location, not the international airport, and data for predictions was collected between 11 and 12 am for consistency in comparison.

Now that summer's over, I have some preliminary results. And the winner (by a hair) is the Weather Network!

Score (out of 100):
  • Weather Network: 66.92
  • Global Weather: 66.02
  • Weather Channel: 63.99
  • Environment Canada: 55.00
  • TimeandDate.com: 54.25
The score is based on a weighted average that was more or less arbitrarily decided by me: each subsequent day in the future was weighted less (so that a prediction for tomorrow's weather is worth more than a prediction for next week's), and POP was worth more than the High prediction, which was in turn weighted more than the Low prediction.

Some fun facts!

Best High temperature prediction: Weather Channel 1-day prediction (96.79% within 3 degrees)
Best Low temperature prediction: Environment Canada 2-day prediction (96.07% within 3 degrees)
Best POP: Global 4-day prediction (p-value 0.346)

Worst High temperature prediction: TimeandDate 6-day prediction (55.20% within 3 degrees)
Worst Low temperature prediction: Global 6-day prediction (68.57% within 3 degrees)
Worst POP: TimeandDate 3-day prediction (p-value 0.038)

Some graphs!

Temperature score was based on the percentage of predictions that were within 3 degrees of the actual temperature. In general there was a very strong downward trend for the high temperature predictions - almost all stations had better than 95% accuracy at predicting tomorrow's weather, and they were all about 70% accurate at the weather a week from now. There was less of a trend noted for the low predictions, however those are typically less useful apart from determining the likelihood of frost.

The score for POP is based off the p-value for each category of prediction. In essence, I checked the number of days that a given station predicted a POP of 10%, and compared it to the fraction of days that it actually did rain for that prediction. This doesn't translate directly into an accuracy percentage, which is why I call them 'scores' instead (though if every category had precisely the incidence of rain as predicted, it would end up with a score of 100).


So there you go! Hopefully this helps you the next time you're planning a picnic (or whatever people check the weather for...).

Saturday, September 8, 2012

Higher Learning


Pick a professional career. Almost any will do. Now really take a moment to visualize how their job is done today - their tools, their projects, all sorts of the stuff that's required by their fields.

Now think about that same profession 300 years ago. Any chance it's changed?

Doctors have progressed from prescribing urine baths and bloodletting to minimally-invasive laparoscopic surgery. Engineers have moved from catapults to Mars rovers, and law enforcement has gone from corrupt court systems to advanced forensics (in most countries, at least). These and other similar advances have been undeniably amazing, and are largely responsible for our currently quality of life.

Sadly, though, one particularly glaring profession has been dragging its heels against this rapid change. Despite the enormous advances over the last hundreds of years, university education has been largely unchanged. For some reason, an overwhelming majority of university classes are still taught by packing vast numbers of students in a theater and being talked to for an hour. Tweed jackets have come and gone, and the use of powerpoint may have sped things up, but by and large the methods used to teach are virtually unchanged.

Courses are still taught largely by talking at students, assigning readings and homework, and then giving grades based on exams. Exams themselves are still mostly just forcing students to cram material for a couple days beforehand, then stuffing students in a room for two hours and making them answer random questions about the previous forty hours of lecture.

Why is this still the standard? How likely is it that, of all professions, how we teach people more or less peaked three hundred years ago? Why is it that the most common way of judging how well someone has learned is to cram them into a room and force them to recite things, and why on earth should the grading that results from that two hour test be worth up to 70% of their grade? The case has been made before that universities should get students focusing on learning how to learn, instead of what often seems to be the focus of getting students learning how to write exams, and I totally agree.

There have been some pretty exciting developments in expanding education options recently, though. For example, the Khan Academy has more than 3,300 video lessons and interactvie exercises covering math all the way from preschool arithmetic to first-year calculus, which are free for anyone to take. The Academy also covers basic sciences, humanities, and finance.

If that isn't what you're looking for, why not learn a language? The BBC offers free courses on all the major European languages, and essential phrases for 40 languages. If you're interested in something less structured, free lecture videos on hundreds of topics can be found anywhere online to anyone who's really interested.

What's particularly cool, though, are the opportunities that are becoming available for a more formalized education. Recently, three of the biggest names in education (MIT, Berkeley, and Harvard) joined together to offer advanced university courses for topics ranging from solid state chemistry to artificial intelligence. These even offer 'certificates of completion' - certainly worth putting on a resumé, even if they don't have quite the same weight as an official transcript. Registration for the courses is open right now, and I strongly suggest you take a look at what's being offered.

The advantage of this new-found variety in fairly high-level education options is that it may (hopefully) end up pushing the envelope for education options in universities. Many post-secondary programs are starting to allow more open-ended education, such as the option to learn by correspondence or online, and with so much knowledge so freely available we may yet see a change in the 'classical' approach to lectures.

And for those of you who truly are here to learn how to learn, I strongly suggest taking a look at some of the links - I can guarantee there's something out there you'll find fascinating.

Monday, August 20, 2012

Scientific Illiteracy

In 1998 British physician Andrew Wakefield was offered £400,000 to publish an article claiming the vaccine against Measles, Mumps, and Rubella could lead to autism. The article was funded by lawyers preparing for an anti-vaccine lawsuit, and despite the fact that his results could not be reproduced and were dismissed by an overwhelming majority of researchers, a massive and ultimately deadly panic swept across the world. Vaccination rates dropped across Europe opening the doors to completely preventable outbreaks, children were killed or permanently disabled, and an estimated $25 million in avoidable hospital bills were accumulated due to completely avoidable MMR outbreaks.

It ultimately took a long and unnecessary campaign by researchers and health professionals to re-prove the safety of the MMR vaccine, but the scientific agreement hasn't quite been translated to the public - in fact today about 48% of Americans either don't trust or are unsure about the safety of vaccines, largely based on that one fraudulent article.

Vaccines are a dramatic example of the consequences of a public that doesn't trust or doesn't understand science, but they are absolutely crucial to public safety. Not getting vaccinated is not only a hazard to yourself but also to everyone around you. Sadly modern outbreaks of completely preventable diseases, such as the recent Whooping Cough outbreak, are reminders of the impact of ignoring science.

Studies have shown, though, that in general the public has a very poor understanding of some fairly basic concepts. Take for example these survey results on scientific literacy:
  • 6% of Americans don't believe smoking can cause lung cancer
  • 13% don't know that plants produce the oxygen we breathe
  • 20% aren't aware the center of the earth is very hot (and are presumably very confused by volcanoes)
  • 25% still think the sun goes around the earth
  • 46% don't know it takes the earth a year to orbit the sun (but were probably still stumped by the previous question), and
  • 52% fell hard for the Flintstones and believe dinosaurs and humans coexisted
Admittedly none of these specific misconceptions of science are likely to be dangerous to an individual, apart from perhaps lung cancer and smoking. What's instead frightening is that these are all concepts that are taught during or before high school, and suggest a public that is largely ignorant or apathetic to some of the most fundamental concepts we rely on. Even more horrifying is that these percentages have tended to only get worse between 2001 and 2010.

Regardless, though, of whether the misunderstanding of science is harmful on an individual basis or not, this attitude towards science of either apathy or automatic distrust is very dangerous for society. Distortion of science for personal or political benefit is very common, and has even been used explicitly to cause harm.

Two particular government exploitations or distortions of scientific understanding come to mind. The eugenics movement in the early 20th century claimed that human breeding needed to be controlled in order to advance our evolution, and was pushed ferociously by political groups and individuals around the world. Even countries like Canada and the United States got caught up in the movement, with individuals like Tommy Douglas and Alexander Graham Bell advocating for restrictions on who could marry and have children, and certain provinces and states forcibly sterilized individuals who were considered unfit to breed. The movement ultimately led to the rationalization of murders in the name of cleansing in Nazi Germany. It ultimately took the combination of the end of World War II and the further understanding of genetics to bring about the end of the vast majority of eugenic based programs, but not after massive personal and societal loss.

On the other hand, Soviet Russia took control of science and dismissed genetics entirely as a "bourgeois pseudoscience", instead adopting the practices of Lysenkoism for agricultural development. This explicit adoption of absolutely useless techniques held back Russian understanding of genetics for decades, and resulted in the firing, imprisonment, and execution of legitimate Russian scientists.

But misunderstanding of science still harms us daily. Despite court findings of fraud, millions of dollars a year of "ion bracelets" are still being sold by pretending to be scientific. Fictionalized versions of polygraph tests have led to the belief that they're foolproof - and private polygraph examiners have likely been responsible for propagating actual lies - even though psychologists have determined that they really aren't any better than guessing. Some people will often reach out to homeopathy at the expense of medically-proven treatments even though it's been consistently debunked by doctors.

Fortunately it's not all bad. The public acknowledgement of Canadian science journalists like Jay Ingram and Bob McDonald with the Order of Canada recently was an important step for supporting the field, which is undeniably important for keeping people informed, and televised outreach through the Discovery Channel and shows like Mythbusters has done a lot for increasing interest in science topics and critical thinking. Hopefully as science continues to advance we will see fewer opportunities for people to take advantage of people's misunderstanding of it, and more opportunities to get people interested and engaged. We definitely need it.

This was also posted on TheWandererOnline with graphics by Michelle Weremczuk. Check it out!

Wednesday, August 8, 2012

Beat the Odds

If you're going to operate a profitable betting ring, there are two things that are important to know. An obvious one is the odds of a particular event happening: a good casino wouldn't make any money if they didn't know the odds associated with blackjack or roulette, and adjusting their payouts accordingly is what gives them any profit.

What's also important to know is the number of people who are likely to bet on a particular outcome. This is less important perhaps on a roulette table, but potentially very important when it comes to something I've been investigating a lot recently: sports betting. It makes a certain amount of sense that sports betting companies (such as SportSelect) track their ticket sales to consumers, and in fact the companies can profit just as much by adjusting their payout according to ticket sales as they can according to individual game odds. As it is likely substantially easier for a company to track their sales than it is to predict the future, this is quite likely a critical factor in their payout odds.

The profit made by a company in sports betting can be visualized like this:


Basically the expected profit can be determined based on the chance of an event occurring multiplied by the cost of that event occurring. So an event (such as home victory) with a chance p of occurring, will pay out an amount equivalent to the payout odds x, and the number of tickets sold that chose that event (m). By considering both at once, we can get an 'expected' average profit that takes into account all possibilities - by tweaking the payout factors, a company can assure themselves of continually profiting from the sports betting.

Presumably SportSelect knows the fraction of tickets sold (m and n) very well. They must. It would be negligence on the part of the company not to know what they're selling, how much they're selling, and who they're selling to. Presumably also they have some sort of model that allows them to predict game odds (p and q) with a reasonable amount of accuracy. As they control the values for payout (x and y), they can then have a good sense of control over their profit.

P and m are both factors that relate to the specific event that's being examined. P is most likely intrinsically involved with the relative strength of the teams, and m accounts for whatever factors lead people to purchase lottery tickets betting on certain teams (perceived skill, popularity, etc.). Together, the factor pm more or less accounts for the expected amount of money SportSelect will lose should that event come to pass.

A quick look at the payout odds that SportSelect offers shows a strong trend - the majority of their payout options average between the two events at a payout of 1.7 - combinations such as 1.6 and 1.8, 1.5 and 1.95, etc. Some more unlikely payout combinations are 1.4 and 2.15, 1.3 and 2.45, and 1.25 and 2.65, and these tend to average a little higher, but still within the range of 1.7-1.95. There are very few combinations outside of that, so for the sake of this piece I'll take only these into account.

In order to try to investigate just how SportSelect comes up with their odds, I set up a series of random p and m values to look at some trends. This is what I got at first:
OH MY GOD THAT'S UGLY. Whew. Jeeze. What I have here is the product of p and m on the x-axis, and then the expected profit percentage based on any of the five payout combinations as explained above (the legend lists the average of the two payout values for each of the five sets) after 1000 data points for each. This is really really ugly though.

Part of the reason it's ugly is the relationship between pm and qn - the complementary payout values for the alternate event. For radical values of either p or m, we tend to get qn values that are tremendously different, which gives the ugly values as shown above. Looking at the cluster of points where the majority lie appears to form a series of curves; this is a cleaner version of that graph:


Much better. This is actually rather interesting, if I may say so myself. What we get is different ranges of the pm factor result in different payout values (x and y from before) giving the largest profits to the company. So for events with either very large differences in who people bet on (m) or who is actually likely to win (p), larger payout odds are more likely to result in profits. Smart, eh? In fact, it's quite easy for SportSelect to guarantee a 14-16% profit by estimating (with not necessarily that much accuracy, even) the pm factor. As they ought to know the sales figures (m), then they only need to be reasonably accurate on the actual game odds in order to make a killing.

Assuming that they do in fact take ticket sales into account, an opportunity to perhaps profit does then exist. Take a look at these tables:

This first one is just a representation of the graph above - each colored zone represents a range where a new payout scheme becomes the most profitable, measured against values for p and m (the middle values are pm). If we change it to represent what those colours actually mean, we get:
In general this follows the trend mentioned before - for games where either the sales or the odds are anticipated to be close, we have lower payouts (an average of 1.7 is typically odds such as 1.6 and 1.8, remember), but with games with larger disparities we have higher payouts offered (an average payout of 1.95 would feature 1.25 and 2.65 for different teams, respectively).

If we look at the pm value with an m of 0.32 and a p of 0.5, for instance, we notice something interesting. The pm value is 0.16, so therefore the most profitable payout distribution would be one with an average of 1.775, such as 1.4 for one team and 2.15 for another. However, if we were really sure that the odds were truly 50-50, then betting on the teams with 2.15 odds against would be profitable - 50% of the time on a $1 ticket we'd get $0, and 50% of the time we'd get $2.15, with an expected return of $1.08. Quickly tabulating these results gives the last graph:

Here the green values are where there's money to be made, the gold values break even, and the red values are guaranteed money losses. These are the values of the absolute best bet that can be made for each combination.

So what does this mean? It means that on certain cases, it could be possible to beat the SportSelect betting system. This would have to involve a very high degree of certainty in the actual odds of a given team winning a game (at least as accurate a model as they use would be required), and it would involve them trying to capitalizing on a fairly significant majority of the public purchasing tickets for one team over another (at least 2:1 ratio would be required).

Still, though, numerically it's possible if you have a good enough model and are patient enough. Good luck!

Tuesday, July 17, 2012

Watch out for thinking traps

Imagine that a new test for HIV has been developed, and that the powers that be have decided to pursue universal testing with it in order to improve public health. Let's assume that it's a very accurate test - it tests positive for 95% of people who actually have it, and it tests negative for 99% of people who don't.

Pretend that you get the blood work done, and get the dreaded bad news of a positive result. Now, you've probably been safe - stayed away from dirty needles, used protection when needed, et cetera, and you immediately go into denial. "This is outrageous," you might say, "it must have been a false positive!"

Taking a cursory look at the information provided about the test, what do you think your chance of a false positive is? Your first impression might be that it's 1%, which certainly seems sadly unlikely.

Let's do the math:

In 2009, the number of people in Canada with HIV was 65,000, and the population was 34 million. That leaves 33,935,000 Canadians who don't have HIV.

The test is positive 95% of the time if someone has the disease, so 61,750 of the people with the disease will get a true positive result. Similarly the test is positive 1% of the time for people without the disease, and would give 339,350 devastating false positives.

That means that of the approximately four hundred thousand Canadians who got a positive result, only sixty thousand of them actually have the disease. If you're not part of a particular risk group, your chance of being healthy with a positive result is not 1%, it's actually a whopping 85%.

This type of approach is critical when dealing with science and technology. A politician could perhaps argue that, with such a seemingly accurate test, everyone should be screened for the good of the public. Due to the low prevalence of the disease, though, such universal testing would do far more damage than good. Actually sitting down and working out the numbers is crucial when dealing with science and statistics, and you should always be wary of blindly listening to information relayed at the beginning, or that's spun against you.

Now that you're primed, let's try a more famous example - the famous Monty Hall problem:

Congratulations! You're in the middle of the game show Let's Make a Deal, and you're standing in front of three doors. Behind one door is a car, and behind the other two are goats. You pick a door, say #1, but it isn't opened. The host, who knows what's behind all three doors, opens one of the other doors, say #2, and shows that it has a goat. The host then says to you, "Do you want to pick door 3?" Is it to your advantage to switch your choice, or does it matter?

Your first reaction may be that it doesn't matter - at this point you are faced with two doors, and you know that one has a goat and that the other has a car. Surely it must be a fifty-fifty decision, and it doesn't matter what you choose. This is, perhaps surprisingly, incorrect. In fact, you can double your chances of winning a car if you switch doors!

Because the host knows what's behind each door, if you switch you are guaranteed to be switching from a goat to a car or vice versa (never a goat to a goat, because that option has been eliminated by the host). Two thirds of the time you're going to pick a goat right off the bat, and switching would get you the car, and only one third of the time would you pick a car first and regret switching. Even though at first glance it would seem your choice doesn't matter, taking a second to think things through can definitely turn out to be profitable to you.

One last example of critical thinking: in front of you are four cards, and you know that each card has a letter on one side and a number on the other. The four cards you can see have E, M, 4, and 8 written on the side facing you, and you can't see the reverse side.

Suppose you were told that the cards always obey the following rule: If a card has an E on one side, then it always has a 4 on the other side. If you wanted to figure out whether this rule is true or false by turning over some of the cards in front of you, which would you turn over to find out? Take a second and think it through. Go on.

Most people immediately choose the card with an E on the front, and rightly so - if it turned out to not have a 4 on the back, we would know the rule is false. Some people will also choose the card with a 4 on the front, but this isn't necessary - the rule was only specified what happened to all cards with an E, not all cards with a 4. Similarly there's no sense in checking the card with the M.

But most people will stop there, and completely overlook checking the card with the 8 on the front. What if you checked it and there was an E on the back? Then the rule would be broken just as easily as if there was no 4 on the back of the card with the E on its front.


This is an example of a confirmation bias: when trying to test theories we often set up tests that are designed to prove something, as opposed to set up tests that are designed to disprove the same theory. Ancient scientists would often accumulate a massive amount of data that agreed with their hypothesis and then call it confirmed, only to be embarrassed when a simple test disproved their theory completely.

Proper critical analysis of probability and the awareness of cognitive biases are important to keep in mind in science and technology, but also in making everyday decisions. The next time you are shocked by something on the news, or need to make a major decision, it always helps to take a step back and really think about  both what you're looking for and the information you've been given.

This post also available at The Wanderer Online.

Tuesday, July 10, 2012

More NHL Streaks

Last week we took a look at the occurrence of streaks in the NHL and found that the rate of winning or losing streaks wasn't any different than what we'd expect from random chance. This was pretty cool in terms of modelling the NHL for next season, as it showed that individual games in the NHL could accurately be modeled as independent of the results of previous games.

A couple people asked me to look into a couple of other possibilities for further investigation. The first thing I decided to look into, and the subject of this post, was the influence of rest time between games on the game outcome. It may make a certain amount of sense that teams playing back-to-back games are at a bit of a disadvantage, for instance, and longer rest times may improve performance. Again, I looked at all the games from the last season for all teams, and this is what I got:

I couldn't think of a great label for the y-axis, so let me describe the graph a little bit better. The x-axis is the number of days since the previous game played by that team. For each category of days since previous game, the percentage of wins was calculated, and then compared to the percentage of games that the team won over all 82 games, which is presented on the y-axis. If the data points are above 1.0, that means that the team did better than average after that number of days off, and vice versa. The grey data points are the results from each team, and the blue squares are the average for each set.

Not so surprisingly, teams tend to do worse when they've only had a day since their last game (on average, they do 20% worse than normal, in fact). Their results tend to improve from there so that they do about average when given a two-day break, and then teams tend to do about 10% better than average when given a three-day break. By the time that a team has had a four-day break, though, they tend to do as well as average again, though with a significantly larger variance between teams.

The results from the 5-, 6-, and 7-day breaks were much less consistent - only about two-thirds of teams had 5- or 7-day breaks, for instance, and only 10 teams had a 6-day break in their schedule. However, on average for the 5+ day breaks, teams tend to do exactly as well as average (no bonus or penalty either way).

Fortunately, the vast majority of games for each team take place after a two-day breaks, and that's fairly consistent between all the teams. Also, the number of one-day breaks is typically about the same as three-day breaks for individual teams, which I'm sure is planned deliberately.

Friday, July 6, 2012

NHL Streaks

Two statisticians get on an airplane to go to a conference. Once they've sat down on the plane, one notices that the other is carrying a bomb with him. When asked about it, the second statistician says, "Well, if the odds of having one bomb on a plane are tiny, then the odds of having two must surely be zero and we're guaranteed to be safe!"

This sort of thinking highlights a very popular gambler's fallacy. Often people will think that if a fair coin has been flipped heads four times in a row, it's more likely to be tails on the next flip because there should be an even number of heads and tails. Though it is true that, after a very large number of flips, we should expect approximately 50% heads and 50% tails, the problem with the fallacy is that each coin toss is completely independent of any other coin toss. No matter how many heads we've flipped, the coin isn't keeping track and trying to balance itself out - we just know that after a while they ought to end up about even.

Statistically independent events are important in probabilities. Coin flips, roulette wheels, dice tosses, and explosive suitcases are typically independent each other, and it's important to know this if you're ever going to go to a casino. (Related: most betting strategies people will try to sell you count on you not knowing this - don't buy them!!). On the other hand, some casino games actually have dependent probabilities. The best example is blackjack - if you know which cards have been played from a deck, then you should know roughly what distribution of cards are left to be played in the deck as it isn't being reset after each hand (which is how card-counting works).

I recently read an article by The Wanderer magazine that examined, in part, a paper by the National Academy of Sciences that investigated the effects of randomness on performance. In the paper, a large-scale model was developed where past performance had an impact on future performance. This got me wondering whether or not this was something I should incorporate into my future NHL models.

Before I get into my results, ask yourself: do you think individual games in the NHL are dependent on previous results? Is a team on a three-game winning streak more likely to win the next game than they would be if they'd just come off a three-game losing streak? Or does it matter?

It's pretty easy to rationalize it either way - perhaps having lost a couple games in a row a team could be feeling depressed and be more likely to lose, or maybe they are more inspired/desperate and would be more likely to win. Similarly, having won a couple games in a row could perhaps make a team more confident and give them an edge, or too cocky and make them lose.

If there is an effect, it would definitely be worth incorporating into a model of the regular season, so I was interested in taking a look at the results of last season. I plotted 2,460 results from the last season (fun), and decided to compare these results to what would be expected if there wasn't any impact from previous games. In every season of 82 games, there are 81 sequences of two consecutive games, 80 sequences of three consecutive games, 79 sequences of four games, etc. On average, we should expect ~50% of two game series to have the same results (win-win or loss-loss), 25% of three games series be a streak, etc. I looked at each sequence of up to 10 games in a row for each team, compared it to what I'd expect without any dependence between games, and this is what I got:

I have to admit I was pretty surprised. The results (red squares), averaged for all teams almost perfectly matched the results that we'd expect if the games were independent. Crazy!

The pink lines on the graph were the maximums and minimums for each streak achieved per team, and the blue lines are the range that we'd expect 90% of teams to fall inside naturally if games truly were independent. Again, most teams fall within this range (oddly enough, typically 27/30 which is our magical 90%).

What does this mean? Basically, for the vast majority of teams, and for the NHL as a whole averaged across all 30 teams, hot streaks or cold streaks from any team throughout the season happen almost precisely as often as we'd expect. There's no evidence here that previous games impact future performance in any significant way.

This is actually pretty good news for my model, because now it doesn't have to be quite as complicated as I was worried about. It's also humbling to know that, even though hockey games depend on the collective actions of a bunch of humans, they follows expected statistical patterns so closely.

So the next time you're betting on hockey games or worried about bombs on airplanes, just take a moment to consider that these events are statistically independent. A team on a hot streak is no more or less likely to win than any other team, and there's really nothing you can do about other people bringing bombs on your flight.