← Finance · Black Swans

Finance · Black Swans

Law of Large Numbers: Can a Billionaire Be 14 Miles Tall?

How many observations do we need to be comfortable that we have a large enough sample size to make a decision? If someone wants to drill down to a core cause of the 2007–2008 financial crisis, they will find it in the incorrect answer(s) to this specific question.

The Law of Large Numbers is a theorem that states the larger the sample size, the closer the sample mean will be to the actual mean of the population. It is also referred to as Bernoulli’s law, after the famous Swiss mathematician. According to LLN, the sample mean (X̄) converges to the true population mean (µ), as the sample size (n) approaches infinity, with probability one. The expected value of a coin toss is .5; there is a 50% chance of heads and a 50% chance of tails, for any single coin toss. As per LLN, if you toss the coin an infinite number of times, you will end up with 50% heads and 50% tails.

The problem lies in two areas: firstly, LLN only holds if the distribution has a finite (determined) mean, that is, α >1 (too technical for this discussion). The second problem lies in the term convergence. We do not have the option to carry out infinite coin tosses, obviously. In fact, our aim of applying this theorem is to get an expected value for the population through the results produced by the (smaller) sample of data available to us. What amount of sample data is enough for the results to converge towards the population mean? How far should our Time series go back in history, in collecting sample data (number of observations), for us to make a comfortable prediction about the future?

Let’s apply LLN on an example to test it out.

Suppose you have been assigned the task to find the average height and net worth of 64 year-old American males. To keep things simple, you decide to settle on a sample size of thirty. You stand in an empty room and, one by one, individuals walk in. You measure their height, ask them their net worth and record the results. After 29 observations, the mean sample height (X-height) is coming out to 5′8″ and the mean sample net worth (X-nw) is coming out to $200,000; about 1.7 inches lower than the true population mean height (µ-height) and around $37,000 below the true population mean net worth (µ-nw). Not bad for such a small sample size.

As the sample size increases, as per LLN, the sample means (of height and net worth) should move (converge) towards the population means, respectively. Just when you are almost done, in walks Bill Gates, as the thirtieth event in the sample. You ask him his height and he replies 5′10″. It moves the average height of your sample up by one inch; bringing it even closer to the true population mean; a change of 1.01%. Then you ask Bill Gates his net worth. He replies, “79 billion dollars.” You look towards your spreadsheet, to enter the data, and repeat, “79 thousand dollars.” He shakes his head and says, “No, 79 billion dollars.” You re-calculate the mean net worth of your sample. It has moved from $200,000 to $2.63 billion; a change of 13,150%. This is equivalent to change in mean sample height from 5′8″ to 14.1 miles. What happened? Why are the two changes in final results so different? Most importantly, what do you do now?

You decide to double your sample size and invite another thirty individuals into the room. The mean height of this new sample of thirty is 5′11″, edging you further towards the population mean. The mean net worth of the new sample of thirty is $300,000. The mean height of the combined larger sample (of sixty individuals) is 5′9.5″. The mean net worth of this larger sample becomes $1.31 billion. Being an astute student of LLN, you start to relax as you observe the Law of Large Numbers in action, moving the sample mean towards the population mean.

However, the sample mean for net worth is still off from the population mean by well over $1.3 billion. You decide to include more 64-year-old American males to your sample. As expected, the mean height keeps moving closer to the population mean by a fraction of an inch at a time. The mean net worth also moves closer to the population mean net worth; however, it, “converges” at a much slower rate than the height. How many 64-year-old American males will you have to bring into the room to get a comfortable value for your mean net worth? Perhaps, the complete population of American 64 year olds may be required. Now, suppose Bill Gates was travelling outside the USA on that day, and never showed up in your sample. What impact would that have on the mean height and on the mean net worth of your sample?

It is through LLN that we reach our second important theorem of probability—the Central Limit Theorem (CLT). It is through CLT that we reach the Gaussian (Normal) distribution. It is through the Gaussian distribution that we reach our models of quantitative finance, econometrics and risk management. If we assume a Gaussian distribution for net worth, then the amount of data required for our sample size can be quite small; anything above thirty seems to suffice. However, in such a (naive) scenario, much like it is impossible for a person to be 14.1 miles tall, it is equally impossible for a person to have a net worth of $79 billion. In probability lingo, a “thin tail” is (incorrectly) assumed for both the physical phenomenon (height) and for the informational phenomenon (net worth) i.e. the same (Gaussian) probability distribution is used to calculate both.

ALL risk management models have made this assumption since times immemorial (1960s to be exact, when quantitative finance took off). This incorrect assumption has resulted in incorrect risk models, incorrect regulatory accords and policies; and subsequently the financial crisis of 2007–2008, when all the risk models across all the banks in the USA (and the world), came crashing down (within, literally, a day)!

So, how much data do we actually need before we can be comfortable with our results in a fat-tailed (informational) world of finance? It is impossible to conclude, theoretically, until we can calculate the existence and impact of the potential Black Swans. Black Swans occur only in the informational domain (stock market, net worth, book sales etc.). A single Black Swan (Bill Gates' net worth, in this case) can alter the value of the mean, overwhelmingly. In statistical lingo, a single tail event dominating the combined impact of all events within the confidence interval. More importantly, while we may be able to gather ALL of the possible data from the past and present (including all historical Black Swans), we still have no mechanism available to model future Black Swans, i.e. just as we convinced ourselves no human could ever have a net worth higher than the net worth of Bill Gates, two days (or months or years) later, someone (let's say Warren Buffett) could walk in with a net worth of $179 billion; while we can say with relative certainty that we will never see a 31.9 mile tall human ((179/79) x 14.1).

The key point to understand is that the speed with which LLN converges for heights of humans (or of any other physical calculation) is much faster than the speed with which it converges for wealth of humans (or of any other informational calculation like the price movements of the stock market). The former results in a Gaussian distribution, while the latter results in a Power law scalable distribution (again, too technical for this discussion).

To put all the above in simple terms: for phenomenon displaying fat tails (like finance), it takes a long (long, long) time (and a lot of data) to know what is going on. If we (naively) assume such phenomenon to be thin-tailed (as we always do in ALL our financial models), we jump too quickly and think we know what is going on; after which we make our risk management decisions, incorrectly. Then we are hit by a Black Swan from the fat tail and we realize (after the hit) that we did not really know what was going on. This is the point where everyone exclaims, “No one saw it coming!”

So, what is the correct answer to the question posed at the start? The correct answer is: "it depends;" on whether we are in the physical (thin-tail) domain or in the informational (fat-tail) domain; are we measuring (the probability of) height or are we measuring (the probability of) net worth (/stock market crash/sales of the post-Harry Potter book series, etc.).

More essays in the Black Swans series are on their way.