Contents
In this chapter, we discuss the relationship that money and debt have with time. The crucial point connecting monetary measures and time is the concept of the interest rate. Think of interest rates as a time machine that allows us to place value on monetary assets at different times by bringing them backward or forward in time to a time when they can be compared. Interest rates are important because they affect: the level of consumer expenditures on durable goods; investment expenditures on plant, equipment, and technology; the way that wealth is redistributed between borrowers and lenders; the prices of such key financial assets as stocks, bonds, and foreign currencies, etc.
For our purposes, we will use “interest rate” and “yield” interchangeably throughout this chapter. We will be more specific in future chapters. For now, we want to begin with some basic calculations about interest rates.
Often, financial instruments yield different payouts at various times. Payments come due at different time periods. And different investments promise returns on different dates. How do we make sense of it all? We must learn how to calculate and compare rates of return on different financial instruments.
Generally speaking, one dollar today is worth more than one dollar at a future date. This is because a dollar today can earn interest and be worth more than one dollar in the future. This will generally hold even if there is no change in the purchasing power of that dollar so long as the interest rate is above zero, which is almost always the case. This also works in reverse. This means that one dollar at a future date is generally worth less than one dollar today.
There are two general concepts that allow us to understand whether the promise to make a payment on one date is more or less valuable than the promise to make it on a different date. Namely, the concepts of future value and present value. We discuss both in turn below.
Future value (FV) is the value on some future date of an investment made today. This is useful to understand the return we will enjoy on an investment at some future date. For example, if we deposited $100 today in a savings account and our bank promised a 5% interest rate if we kept the $100 in the account for one year, at the end of the year we would have $105.
How do we come up with that number? We would get back our original $100 amount of savings. But we would also collect a return from having saved the amount at an interest rate of 5%. The interest return is $5 since .
Therefore, $100 invested today at 5% interest gives $105 in a year. We can say that the future value of $100 today at 5% interest is $105 one year from now. The initial $100 yields $5, which is why interest rates are sometimes called yields. This scheme is like a simple loan of $100 for a year at an interest rate of 5%. So, in the present time, the value of this investment is $100, which in the future has a value of $105. Altogether, this is calculated as . This gives us some insight into our first equation:
where the future value (FV) of the present value (PV) of $100 investment—when the interest rate (i) promised for the next year is 5%—equals $105.
What if we left the $105 for another year with the interest remaining unchanged at 5%? This question can be restated as: “What is the FV one year from now of $105 at present at an interest rate of 5%?”
Applying the formula shows we would have $110.25 next year. . This was the answer to our decision to reinvest the $105 once we collected it from our initial $100 investment the previous year.
What if we had made that plan from the beginning? If we wanted to know how much would we have in two years if we invest $100 today at an interest rate of 5%?
In two years, we would get our $100 back, plus the interest earned in the first year , plus the interest earned on the investment the second year (another ). But after year one we would have collected $5 in interest earnings. That interest earning would also collect interest on the second year (this type of interest accruing interest is called compound interest), which would amount to .
Another way to ask the question would be: What will the future value of $100 be in two years when the interest rate is kept fixed at 5%? The future value would be , the same answer we had before. But this could easily be rewritten as .
So while the equation above holds for the FV if we only invest for one year, if we invested it for two years, the formula becomes . So, in general, the future value (FV) of a current investment at an arbitrary number of years in the future (n) at a fixed interest rate (i) is given by:
gives the future values of $100 investment at a fixed interest rate of 4% for various years into the future.
Table 3.1Future value calculations.
| Years into the future | Calculation | Future value |
| 1 | ||
| 2 | ||
| 3 | ||
| 5 | ||
| 10 |
Interest earned on the original principal is called simple interest (SI). Interest earned on interests already paid is called compound interest (CI). Total interest is the sum of the two. In the example above, total interests are $10.25 after two years. In that example, and .
We notice SI is much larger in this case ($10) and CI is much smaller ($0.25). But this is because we invest over a small number of periods. Investing or saving in the short term, we find SI is always larger than CI. However, while SI grows linearly, CI grows exponentially. Therefore, the power of compounding clearly becomes more important as time passes.
To make this point clear, take the following example. On May 24, 1626, Peter Minuit purchased Manhattan Island from the Canarse Native Americans for about $24 worth of trinkets, beads, and knives. The purchase took place at what is now Inwood Hill Park in upper Manhattan, New York. Consider the following thought experiment. If the tribe had taken the cash value, instead of the goods, and invested it to earn 6% per year, how much would the tribe have in present day, about 387 years later?
This can be easily calculated to be roughly $267 bn with a FV equation in Python.

The total interest that would have been earned is, in essence, the FV of the investment scheme minus the original $24 invested, which in Python code can be expressed as:

But how much of this total interest is earned simply from the original amount invested? Will it be more or less than the interest earning interest? In other words, how much of this is gained from SI and how much from CI?
The amount of SI 387 years later is a miserly $572.

The rest of the $267 bn is earned with CI!
The assumption we have been making so far is that interest rates are paid/earned once a year. What if interest rates were paid monthly? Converting n from years to months is easy, but converting the interest rate is a bit harder. If the annual interest rate is 5%, what is the monthly rate? Assume is the one-month interest rate and n is the number of months, then a deposit made for one year () will have a future value of:
Since we know that in one year the future value is $100(1.05), we can solve for :
Therefore , or 0.0041. These fractions of a percentage point are called basis points (or bps). One basis point is one one-hundredth of a percentage point, 0.01 percent. On a monthly basis, the FV of $100 one year from now at an annual rate of 5% would be 41 basis points.
Another common question for an investment proposition would be: If we invest $100 at 5% annual interest, how long will it take to double our money? A rough approximation can be obtained with what is known as the Rule of 72. We divide the annual interest rate (in percentage points, not decimals) into 72.
If the interest rate remained fixed at i%, the approximate time to double the investment is 72/i, independently of the amount.
In our example, the interest rate was 5% so it would take (72/5 =) 14.4 years to approximately double our investment.
To verify this, suppose we save $5,000 today at an interest rate of 5%. After 14.4 years, we would have
essentially doubling our initial investment.
The present value (PV) is the opposite side of the coin of our concept of future value. And it is quite useful in finance because PV is like a time machine that allows us to place a value on payments that will be received in the future in today’s dollars. In fact, according to the classical theory of asset prices, the price of a financial or monetary asset equals the present value of the expected asset income that it promises in the future. Asset prices depend on people’s expectations of future asset income, and PV is the tool we use to compare the desirability of different investment schemes, even if they only promise payments in the future and do not deliver anything now.
Since many financial instruments, pension schemes, and investment opportunities promise future cash payments, we need to know how to value those payments without having to wait until they are actually made.
Present value (PV) is the value today (in the present) of a payment that is promised to be made in the future. Or, alternatively, it is the amount that must be invested today in order to realize a specific amount on a given future date. Mathematically, we can simply solve the FV formula for PV as follows:
which is just the future value calculation inverted. We can generalize the process for arbitrary n periods in the future, as we did for the future value, to answer the following question:
What is the PV of a future payment made n years in the future, at a fixed interest rate i?
This can be answered with the following equation:
We can conclude a few things from this equation. First, the higher future value of the payment, the higher its present value.
Second, the longer into the future we have to wait for a future payment to be made, the less valuable it is today. That is, the PV decreases as n (the time to payment) increases. This is because the longer we have to wait for the payment, the longer we are effectively discounting it.
Third, the present value of a future payment goes down as the interest rate increases. Think of the interest rate as the reward we would be getting for holding an alternative asset that promises it. The higher the interest rate, the more interest return we are forgoing had we invested in the alternative asset that pays it. Therefore, the higher the interest rate, the faster we discount the future value. This lowers the present value.
Doubling the future value of the payment, without changing the time of the payment or the interest rate, doubles the present value. This is true for any percentage.
Table 3.2Present value calculations.
| Interest rate | 1 year | 5 years | 10 years | 20 years |
| 1% | ||||
| 2% | ||||
| 3% | ||||
| 4% | ||||
| 5% | ||||
| 6% | ||||
| 7% | ||||
| 8% | ||||
| 9% | ||||
| 10% | ||||
| 15% |
shows the PV of a $1,000 payment. Higher interest rates are associated with lower present values, regardless of size or time of payment. As we select a given interest rate in a given row, and we travel rightward across the columns, we can also see that at any fixed interest rate, an increase in the time reduces its present value. The longer we have to wait, the more we discount the future payment, and the higher the interest rate, the faster we discount the future payment.
Present value is the single most important relationship in our study of financial instruments. One reason PV calculations are so useful is because they allow us to compare different payment options in time.
For example, Annie just won a lottery worth $5 million. The lottery will pay her over a period of 5 years (i. e., $1 million per year for 5 years). How much did she really win today?
In order to answer that question, she needs one piece of crucial information: what is the prevailing interest rate in the market?
This will be important for Annie to make her decision. If the interest is high, she would be missing out on interest gains while she waits to get paid, because she cannot collect interest on money she does not yet have. This means waiting longer is more costly when interest rates are high. If the interest rate is low, future payments are discounted more slowly, so she might prefer to wait as the opportunity cost she is faced with is low when interest rates are low.
Suppose that the interest rate is 2% and she is given two choices: take $1 million every year for the next 5 years (this is an installment plan) or take $4 million now (this is a lump-sum payment). How could she compare the two options? She knows the PV of $4 million today is $4 million. What is the PV of the installment plan?
She needs to compare the $4 million payoff with the PV of the other option of the installment payment. Whichever plan offers her the highest PV should be the more desirable option.
The installment plan involves the same cash flow of $1 million at an interest rate of 2%. The first one must be discounted one period (), the second payment in two years must be discounted two periods (), and so on. We must apply our PV formula for each annual payment and then add them all up. The following bit of Python code will reveal the PV of the proposed installment is $4.7 million. Therefore, the installment plan is worth more than the lump sum payment.

But we are missing one last piece of information: the prevailing interest rate in the market. We can code it into Python as follows:


The following code defines a PV equation that calculates the present value of this plan is $4,713,459.51:

In this case, at an interest rate of 2%, the present value of the five-year payment stream is $4.7 million, which is greater than the $4 million lump-sum payment so Annie might take the annual payment option.
But what if the interest rate was substantially higher, say 9%?
This means that future payments would be discounted faster. If Annie got the money now, she could deposit it in a bank account that would pay her 9%. Not having the money now to deposit is costing her a potential 9% return. This means she might be less patient and less willing to wait for future payments if the interest rate is this high. So the question is: Should she take the $1 million annual payment for five years at a fixed interest rate of 9% or the lump-sum payment of $4 million today?
In Python code, we can easily compute the PV of this payment as:

which comes out to be $3,992,710.04. Now, the present value of the payment stream is less than the lump-sum payment, so Annie would probably take the one-time (lump-sum) payment.
But this should make sense. A much higher interest rate means the opportunity cost of holding cash right now is higher. Having to wait to get paid means Annie would be forgoing a higher return on her cash. Therefore, the value of the annuity is discounted faster.
Consider a second example. Let’s say Annie has a cousin named Paul. Paul wants to borrow $10,000 from Annie today and promises to pay her back in 20 years at an interest rate of 6%. This means that the future value of the loan—what Paul will need to pay her back in 20 years—is $32,071.36. Both Annie and Paul can make this FV calculation. They both agree that waiting 20 years is too long to pay off the loan and choose an installment plan.
Paul proposes dividing the FV equally among the 20 years to make an annual installment payment every year for the next 20 years. While the proposal may sound fair, Annie wants to know if Paul is correctly valuating the present value or whether he would be overpaying or underpaying what is owed. They both know the PV of the $32,071.36 loan in 20 years time is $10,000 today. The question is: What is the PV of Paul’s proposed installment plan? In the previous example, we knew the payment amounts, the number of years, and the interest rate, and we wanted to calculate the PV.
In this example, we know the PV, the number of years, and the interest rate, and what we want to know now is the actual payment amount—the cash (C) disbursement to be made every year. We apply the PV formula as follows:
Annie and Paul want to know what C will be. Will it really be $1,603.57? This can be easily calculated in a few lines of Python code.
Instead of writing the PV equation every time and redefining our values (, n, i) each time, we could automate our analysis by defining a function that calculates the PV when supplied with inputs.
The snippet of code below defines a function called PV_simple to calculate the PV of a single cash flow (CF).

The code above defines the PV function. Once that function is instantiated by entering into memory in Python, it can be called simply by entering values for , i, and n. We can then call the function with the appropriate values to verify the PV of the loan is $10,000.

Now we need to calculate the PV of the installment plan. We need a different equation than PV_simple. We first define an equation to calculate the PV of an installment plan paid yearly, which is generally termed an annuity.

This is a very similar equation to PV_simple, but instead of taking a single value of a cash flow, it will take multiple values (a range) of cash flows. Paul’s proposed plan has 20 payments of $1,603.57. The following lines of code instantiate the payment plan:

which outputs the following list of values:

We can convert a list of values to a numpy array, which makes it easier to work with as follows:

which outputs the following:

Applying this 20-year installment plan and the relevant information to the PV_annuity function we defined earlier allows us to find out that at the PV of the proposed payment plan is almost twice as much as the that is owed.
The PV of the loan is $10,000 yet the PV of the proposed payment is $18,393. This means each year Paul would be overpaying almost twice what the correct amount should be. What would be the correct annual payment to guarantee a $10,000 PV under these terms?
Again, the correct payment (C) solves the following equation:
which in Python, can be solved with:

which shows the correct payment of this installment plan should be . At , the payment is roughly half what Paul proposed. Annie wants to verify this by calculating the PV of 20 annual payments of $871.85 at an interest rate of 6%.
Applying our defined formula, she can verify the PV is indeed up to a rounding error.

This is an example of a fixed payment loan. However, all this analysis can extend to a variable payment loan where the cash flow (C) changes over time. This is important because many securities are claims on future payments, which may vary. This allows us to employ the PV calculation to evaluate prices of securities. In general, the price of any security can be obtained as the present value at a given interest rate or yield of the future payments expected to be made by the security issuer. The same equation can be used where the payments, instead of being fixed cash flows (C), might be variable payments (R), as follows:
Thus far we have assumed interest rates (i) are fixed. But the fact that future payments may vary over time introduces some uncertainty as to the amount of a future payment or whether the payment will be made at all. This introduces some measure of risk. In the face of uncertainty, investors will generally want to be compensated for taking on risk. This is called a risk premium. This premium typically takes the form of a yield that is added above and beyond the prevailing interest rate (i). A risk-adjusted PV equation takes the following form:
where ψ represents what is commonly known as the risk premium. As investors demand higher compensation for taking the risk associated with the uncertainty of a future return that is not guaranteed, the value of that future payment is further discounted above the prevailing interest rate. In other words, as risk of a future payment becomes riskier, investors demand a higher risk premium (a higher ψ), reducing the present value of the investment lower than it would have been had the future payment been guaranteed.
Thinking About it…
Interest rates act as a “time machine” for comparing monetary values across different time periods. The time value techniques of PV and FV, which we cover in this chapter, can be used to price financial instruments and securities.
FV shows what an investment today will be worth at a future date. It incorporates both simple interest (earned on principal) and compound interest (earned on previously earned interest). It demonstrates the power of compound interest through examples, like how $24 invested in 1626 at 6% would be worth billions today.
PV can be used to compare different payment options (like lump sum vs. installment payments). For a given payment scheme, PV decreases as interest rate increases (faster rate of discounting), or when the time to payment increases (longer period of discounting).
PV and FV are useful because securities differ in their future payments. For example, bonds generally have fixed payments and U. S. government bonds have no default risk, while corporate bonds are often subject to default risk. On the other hand, stocks have uncertain payments and they are also subject to default risk.
The bigger the risk premium or the higher the yield or interest rate, the lower the price of a security. In order to understand risk, we must quantify it. This is done with the mathematics of probability theory, which we cover in the next chapter.
The process of gradually paying off a debt over time through regular payments that include both principal and interest. Each payment reduces the principal amount owed.
A series of equal payments made at regular intervals over a specified period, often used in loans, mortgages, or retirement planning.
A unit of measurement for interest rates, where one basis point equals 0.01% (one one-hundredth of a percentage point). Used to precisely describe small changes in interest rates.
The profit realized from the sale of an asset when its selling price exceeds its purchase price or original investment.
A theory in the field of finance that posits that the price of a financial or monetary asset equals the present value of the expected asset income that it promises in the future.
Interest earned on previously earned interest, rather than just the principal. While smaller than simple interest in the short term, it grows exponentially over time and can lead to significant wealth accumulation over long periods.
The possibility that a borrower will be unable to make required payments on their debt obligations. This risk often influences the interest rate charged on loans.
The interest rate used to determine the present value of future cash flows. Often reflects both the time value of money and risk considerations.
The value of an investment at a future date, calculated by accounting for the initial amount (principal) plus any accumulated interest over time.
A measure that acts as a financial “time machine,” allowing comparison of monetary values across different time periods. It represents the rate at which money grows when invested or the cost of borrowing money.
The difference between the present value of cash inflows and outflows over a period of time. Used to analyze the profitability of an investment or project.
Numpy is a Python library that is used for numerical calculations. A numpy array is a type of data formatted in Python that can be organized in vector or array form, which facilitates calculations.
The current value of a future payment or series of payments, discounted back to the present using an interest rate.
The original amount of money invested or borrowed, before interest begins accruing.
The return on an investment after adjusting for inflation, representing the actual increase in purchasing power.
An additional return demanded by investors above the basic interest rate to compensate for uncertainty in future payments. The higher the risk, the higher the premium required, which results in a lower present value of future payments. This compensation investors demand for taking on added risk is typically expressed as a percentage.
A quick method to estimate how long it will take for an investment to double at a given interest rate. Divide 72 by the interest rate percentage to get the approximate number of years.
Interest earned solely on the original principal amount invested. It grows linearly over time and is typically larger than compound interest in the short term but is overtaken by compound interest over longer periods.
The concept that money available now is worth more than the same amount in the future due to its potential earning capacity through interest.
Used interchangeably with interest rate in basic calculations, representing the return on an investment over a specific period.
Contents
Risk is ubiquitous in financial markets. Future outcomes are uncertain. Risk may be unavoidable but, as it turns out, possibly useful. In order to mitigate risk and possibly gain an advantage over uncertainty, we must attempt to assess it first. Assessing risk involves finding ways to quantify it. In this chapter, we consider a brief primer on the concepts of probability. Probability allows us to quantify the likelihood of future outcomes. These concepts can then be leveraged to quantify risk in financial markets.
There are three main axioms of working with probabilities that are attributed to the Soviet mathematician Andrey Kolmogorov. The three main Kolmogorov axioms are:
Axiom 1: Probability can be represented as a real number greater than or equal to zero and less than or equal to one.
Axiom 2: Total probability of some event is equal to 1.
Axiom 3: Probability of mutually exclusive events is the sum of probabilities.
Probability can be thought of as an area. We use to denote the probability of event A, or the area of A. The first axiom suggests . Zero means A occupies no area and 1 means A occupies the whole area.
Probabilities of events can also be added up. For example, represents the probability that event A and event B occur.
The probability of an event can also be thought of as the frequency with which that event occurs. Take a fair-sided dice, which means each side is equally likely, with six possibilities. If x denotes a given side of the dice, say 1, then its probability is given by . This is an example of discrete probability, because each roll of the dice is a discrete and independent event.
A continuous probability refers to a process where events are not easily separable at discrete intervals. One way to generate a continuous probability is to repeat a single discrete event (e. g., throwing a dart) many many times, so that the event is so frequent that it is difficult to separate any given outcome from the other. Throwing 10 darts at a board could be modeled as a discrete probability, whereas throwing 10 million darts could be more easily thought of a continuous probability.
Let’s say we want to throw a dart at a circular target inside a square with two meters on each side. Each throw is an independent event. See .

Figure 4.1 A target for many dart throws.
What would be the probability of hitting the orange target C? The area of C is π, and the area of the square that contains it is 4 (2 m × 2 m). Now, if we divided the target in four quadrants, what would be the probability of hitting the top right (northeast) quadrant?
The probability of hitting the orange target in that upper right quadrant should be . We can verify this with a simple simulation. First, we note that each dart throw is an independent event. We are going to simulate it as a random event. The following snippet of Python code simulates 1,000 throws. It draws the portion of the throws that would land inside and outside the target and graphs them.

This Python snippet carries out a simulation that generates 1,000 throws out of which 796 land in the orange target.
looks somewhat sparse given we have only 1,000 simulated values. But we could really generate a much larger simulation. If we do this again for 100,000 throws, we get something closer to the theoretical value of . The following snippet shows a similar figure, now filled in with many more simulated dart throws. All we need to change is to substitute 1K for 100K in the first line. See .


Figure 4.2 A Python simulation of 1,000 dart throws.

Figure 4.3 A Python simulation of 100,000 dart throws.
The following command

demonstrates that our simulation provides a rough approximation to the measurement of . This technique is known as Monte Carlo integration and is very commonly used when we are interested in estimating values of complex functions.
Comparing the two simulations reveals the 1,000 throw simulation yields , whereas the 100,000 throw simulation shows , which is closer to the theoretical value of . Increasing the number of simulated throws provides a better approximation.
What if instead of a single event, we have a sequence of events? If the events happen in sequence, but are otherwise independent of each other, we simply multiply the probabilities. For example, in a fair-sided dice, if the probability of getting a three is , then the probability of first getting a three and then getting a two is:
Similarly, the probability of getting an odd value on the first roll and an even value on the second roll would be:
which also works in reverse order of throws:
The general methodology involves three steps: first, we enumerate all possible outcomes; second, we calculate the “area” of each outcome; and third, the probability of each outcome is calculated as the fraction of the total area it occupies. To enumerate all possible outcomes, it is useful to borrow ideas from combinatorics.
Permutations are the total number of possible sequences of n different elements, which are given by a concept that is known as “factorial”, which is denoted by the ! sign, where means:
Another concept of combinatorics is Combinations. They show the total number of ways of grouping N elements into two groups of size k and .
With these two concepts, we can easily calculate the number of possible outcomes in many situations. For example, the number of possible playing card shuffles (possible permutations) is in scientific notation, a very large number. Another example is the number of ways of getting three tails and two heads in five coin flips (combining three tails and two heads in five flips) is given by
There are three useful tools to explain probabilities:
Histograms: Summarize the observed frequency of each outcome .
Probability Distributions: Show the probability associated with each possible outcome .
Cumulative Probability Distributions: Show the probability associated with outcomes smaller or equal to each outcome .
Can we calculate the probability distribution of the outcome of flipping five coins that come out heads with probability p?
The probability of flipping heads and tails is given by , since we need to get heads with probability p each time and tails with probability each time. As we saw earlier from combinatorics, we also know the number of ways of getting heads out of a total N flips is given by:
Therefore, combining the two equations allows us to compute the probability distribution of getting heads out of a total N as:
The following one line of Python code applies this equation to the probability of three heads in five throws.

Let’s simulate the probability distribution in Python. The snippet below defines a simple function to generate coin flips where , . We can specify how many coins we want to flip at each step and how many steps we want and what the probability of flipping heads is.

We can call this function to flip a single coin five times and assign it to a variable called “coins” as follows:

We can always turn the numerical value to a label heads or tails with the following snippet of code:

A histogram is a tool that summarizes the distribution of numerical data. The term was first introduced by the statistician Karl Pearson. To construct a histogram, the first step is to divide the entire range of values into a series of intervals—and then count how many values fall into each interval, called bins, which are placed next to each other.
A probability mass function is a tool that can be used to calculate histograms. The following snippet of Python code defines a probability mass function:

Let’s say we want to figure out the number of ways to get three heads and two tails by flipping five coins. We would need to flip five coins at once, record the result and repeat it again. Once we toss the five coins a large number of times, we may become more assured we have covered all possible combinations. We can do this in Python by simulating flipping five dice 10,000 times and construct the probability distribution of the simulation and compare it to the theoretical probability distribution with a histogram.

These next lines produce the probability distribution of the 10,000 simulated flips and calculate the theoretical probability of getting three heads and two tails (or two tails and three heads) in five flips, as well as producing graphs of the results:

By graphing both, we can conclude that the simulation approximates the exact distribution closely. See .

Figure 4.4 Probability distributions of three heads and two tails in five flips.
Naturally, we expect the curve to be symmetric as there are exactly as many ways of having three heads and two tails as there are of having three tails and two heads. There are many different types of probability distributions—some of which may be symmetric and some of which may not be.
An n-sided dice has a uniform probability of landing on any of its sides. After, say, 10 rolls of a fair dice, we might have: . The behavior of this x variable is stochastic, which is another word for probabilistic. But what about a function of this random variable (like the average)? The equation for the average is given by:
In this specific example, the average is 3.5, as expected. But if, say, the sixth and the seventh rolls had come out sixes instead of ones, the average would have been 4.5.
We could throw the dice 10,000 times and we could find an average. But if we repeated the experiment 10 times, we would find 10 different averages, so what would be the correct value of the average?
For any given set of dice rolls, we are only approximately estimating the true value. In general, the higher the number of rolls in a given realization, the better our estimate of the average. The true, or expected value, is the one obtained after an infinite (or at least, very large) number of realizations.
The “true average” or the expected value of a random variable involves multiplying each outcome times the probability of its occurrence and adding them up as follows:
The central limit theorem (CLT) is a useful result in probability theory. It allows summarizing large samples into a sample average, because they can approximate the true expected value if the sample is large enough (if it has enough observations). The CLT advances that for independent and identically distributed random variables, the standardized sample average tends toward the standard normal distribution so long as the number of observations is large enough. Even if the original variables themselves are not normally distributed, their sample average can be approximated by a normal (also known as a Gaussian) distribution.
Let’s say we obtain a large sample of observations that are independent from each other—each generated observation does not depend on the values of the other generated observations—and we compute their sample average (or arithmetic mean). If we repeat this procedure many times, resulting in a collection of observed averages, the CLT says that if the sample size was large enough, the probability distribution of these averages will closely approximate a normal distribution.
The probability of observing a random value of x from N observations drawn from a normal distribution centered at its mean and with variance is given by:
The mean value is given by and the variance is . If the sample size N is large enough, it can be approximated by a normal distribution.
The CLT is a key concept in probability theory because it implies that probabilistic and statistical methods that work for normal distributions can be applicable to many problems involving other types of distributions.
Thinking About It…
Imagine we are standing in a trading pit in a financial exchange like the New York Stock Exchange (NYSE), watching investors buy and sell securities. Each of those securities have uncertain returns. How can traders price those securities if their returns are not fixed and guaranteed? A way to do this is to think probabilistically, starting with the fundamental rules laid down by mathematician Andrey Kolmogorov.
Kolmogorov teaches us that we can think of probability like measuring the area of possibility—it can’t be negative, and it can’t exceed the whole space of what’s possible.
We can think about simple and countable events (like throwing a single dart on a dartboard or buying a single share of stock once). However, financial markets are rarely built on a single transaction, so it requires the complex world of continuous probability. Imagine throwing one dart at a target—that’s discrete. Now imagine throwing ten million darts—suddenly the individual throws blur together into a continuous pattern. Combinatorics help parse out large amounts or flows of transactions.
A single transaction may look random. And when combining it with may other random transactions, the complexity may look chaotic. However, the remarkable central limit theorem (CLT) principle tells us that when we look at enough random events together—whether they are dart throws, stock prices, or interest rates—their average behaviors tend to follow a predictable pattern. Probability and statistics can be leveraged for coping with the inherent uncertainty of financial markets.
A fundamental principle or self-evident truth on which other statements are based. In probability theory, there are three main Kolmogorov axioms that form the foundation of probability theory.
A fundamental principle stating that for independent, identically distributed random variables, their standardized sample average tends toward a normal distribution as the sample size grows larger.
Probability distribution where events are not easily separable into distinct intervals. Emerges when discrete events are repeated so frequently that individual outcomes blend together (like millions of dart throws).
A statistical measure that expresses the extent to which two variables are linearly related, ranging from −1 to +1.
A function that shows the probability associated with outcomes smaller than or equal to each possible outcome.
Probability that deals with separate, countable events where each outcome is distinct and independent (like dart throws or coin flips). Each event has a specific, separate probability.
The “true average” of a random variable, calculated by multiplying each possible outcome by its probability and summing these products. Represents the long-term average after infinite trials.
The three fundamental principles of probability theory: probabilities are non-negative numbers between zero and one, total probability equals 1, and the probability of mutually exclusive events is the sum of their individual probabilities.
A theorem stating that as the number of trials increases, the sample average tends to converge to the expected value.
A technique using random sampling to obtain numerical approximations of complex mathematical functions and probabilities.
Events that cannot occur simultaneously, where the occurrence of one event prevents the occurrence of the other.
A symmetric, bell-shaped probability distribution defined by its mean and variance, which often emerges as the limiting distribution in real-world phenomena explained by the central limit theorem.
A statistical hypothesis that assumes no significant difference exists in a set of given observations.
The complete set of all items or individuals that are of interest for a particular statistical study.
A mathematical description showing the probability associated with each possible outcome of an event. It can be visualized as a graph or formula.
A numerical measure between zero and one that represents the likelihood of an event occurring. Can be thought of as the “area” of possibility that an event occupies within the total space of possible outcomes.
The arithmetic mean of a set of observations, calculated by summing all values and dividing by the number of observations.
The number of observations or data points in a statistical sample, important in determining the reliability of statistical estimates and the applicability of the central limit theorem.
Another term for probabilistic, referring to random variables or processes that can be analyzed statistically but not predicted precisely.
A probability distribution where all possible outcomes have equal likelihood of occurring, such as in a fair die.
A measure of variability in a dataset, calculated as the average of squared deviations from the mean.
Contents
According to the Merriam-Webster Dictionary, risk is “the possibility of loss or injury.” For outcomes of financial and economic decisions, we need a different definition. Risk is a measure of uncertainty about an investment’s future payoff. Risk is a relative concept. It is always assessed over some time horizon and compared to a benchmark. In this chapter, we familiarize ourselves with the concept of risk and the mathematics necessary to quantify risk.
Risk is an abstract concept, but it can be quantified. Generally speaking, the riskier the investment, the less desirable and the lower the price. Risk arises from uncertainty about the future. We do not know which of many possible outcomes (economists call this “states of the world”) will follow in the future. Because risk has to do with the future payoff of an investment, we must imagine all the possible payoffs and the likelihood of each.
The definition of risk refers to uncertainty over an investment or group of investments. Since what constitutes an investment can be broad and not very specific, the concept of risk is often described very broadly. However, there are two main tenets that always apply to risk. One, risk must be assessed over some time horizon. In general, risk over longer periods is higher. Two, risk must be measured relative to some benchmark—not in isolation. A good benchmark can be the performance of a group of experienced investment advisors, money managers, or even the market index itself.
In order to quantify risk, we need to familiarize ourselves with the mathematical concepts surrounding random events and probabilities. One important concept borrowed from probability theory is that of expected value, which summarizes a given investment’s return out of all possible values.
To understand investment returns, we need to list all the possible outcomes and figure out the chance of each one occurring.
Probability is a measure of the likelihood that an event will occur. It is always between zero and one. It can also be stated in terms of frequencies.
Assume we have an investment that can rise or fall in value. Let’s say that for a $1,000 investment in the stock market, the value of our stock can rise to $1,400 or fall to $700. The amount we could get back is called the investment’s payoff. We can construct a payoff table and determine the investment’s expected value—the average or most likely outcome.
Table 5.1A payoff table for investing $1,000 (Case 1).
| Possibilities | Probability | Payoff | Payoff × Probability |
| #1 | 0.5 | $700 | |
| #2 | 0.5 | $1,400 |
Expected Value = Sum of Probabilities × Payoffs = $1,050.
shows a probability payoff for investing $1,000 based on two states of the world with equal (0.5) probability. This table shows that a $1,000 investment has an expected value of $1,050. This does not mean the investment guarantees a payoff of $1,050. Indeed, if we invested only once, we could never get $1,050 back; we would either get $700 or $1,400. The expected value shows the average of most likely outcome of the investment if we were to repeat it many times.
Repeating the investment multiple times, sometimes we would receive the $700 payoff and sometimes we would receive the $1,400 payoff, which would average to $1,050—particularly if we ran the investment enough times for the central limit theorem (CLT) to apply.
Let us now imagine we have a different investment scheme where a $1,000 investment may lead to more possible outcomes. Let’s say our $1,000 investment could rise in value to $2,000, with a probability of 0.1, or could rise in value to $1,400, with a probability of 0.4, or could fall in value to $700, with a probability of 0.4, or could fall in value to $100, with a probability of 0.1. This information can be summarized in the following payoff table.
Table 5.2A payoff table for investing $1,000 (Case 2).
| Possibilities | Probability | Payoff | Payoff × Probability |
| #1 | 0.1 | ||
| #2 | 0.4 | ||
| #3 | 0.4 | $1,400 | |
| #4 | 0.1 | $2,000 |
Expected Value = Sum of Probabilities × Payoffs = $1,050.
shows a probability payoff for investing $1,000 based on four states of the world with varying probabilities. Multiplying the payoff times its corresponding probability and summing these across all possibilities reveals the expected value of this second investment scheme is also $1,050. Using percentages allows comparison of returns regardless of the size of initial investment. The expected return in both cases is $50 on a $1,000 investment, or 5%.
If both investment strategies (Case 1 and Case 2) yield the same return of 5% on an initial investment of $1,000, does this mean the two investments are equivalent?
Not necessarily. We note that Case 1 has only two possible outcomes, whereas Case 2 has more outcomes. The possibility of more outcomes gives rise to more uncertainty on what return an investment will yield. This means that while the expected value gives important information about the desirability of a given investment strategy, we must turn to another concept to measure the uncertainty or risk associated with a given investment strategy.
It seems intuitive that the wider the range of outcomes, the greater the risk. A risk-free asset is an investment whose future value is known with certainty (a single state of the world and a single outcome with a probability of one) and whose return is the risk-free rate of return. This means the received payoff is guaranteed and cannot vary. Neither Case 1 nor Case 2 above are free of risk, because the payoff can vary across states of the world (two for Case 1 and four for Case 2).
But their expected values do not give us an insight on how risky these two investment strategies are. To quantify their risk, we must come up with a measure of their spread. Measuring the spread in payoffs allows us to measure investment risk.
One measure of spread is given by the variance. The variance is the average of the squared deviations of all the possible outcomes from their expected value, weighted by their probabilities. This means the variance can never be a negative number. The smallest value a variance can take is zero. An investment with a variance of zero is an investment with a constant payoff (one that never varies). This means the payment is guaranteed and it is, therefore, an investment with a risk-free return. As the variance rises above zero, the risk of investment increases. The higher the variance, the higher the risk.
An equation for the variance is:
Therefore, the steps to compute the variance are:
Compute the expected value: We can compute the expected value as: .
Subtract expected value from each of the possible payoffs and square the result: (dollars)2; (dollars)2.
Multiply each result times the probability and add up the results: .
This gives us that the variance of the investment for Case 1 is not zero, so this investment scheme is not free of risk. The question is whether this variance is large or small relative to alternative investment schemes. One issue with the concept of variance is that the units of measurement may be difficult to interpret. We do not know what the square of a dollar really is! So an alternative measure is called standard deviation, which is simply the square root of the variance. The standard deviation is more useful because it deals in normal units, not squared units (like dollars-squared). Also, we can calculate standard deviations into a percentage of the initial investment, so it can be more readily interpreted.
The standard deviation is given by:
which in Case 1 is simply . This means the standard deviation of the initial investment of $1000, is $350 or 35%. Let us now calculate the standard deviation of Case 2. See .
Table 5.3Calculating the standard deviation of investment $1,000 (Case 2).
| Probability | Payoff | Payoff-expected value | (Payoff-expected value)2 |
| 0.1 | $902,500 (dollars)2 | ||
| 0.4 | $122,500 (dollars)2 | ||
| 0.4 | $1,400 | $122,500 (dollars)2 | |
| 0.1 | $2,000 | $902,500 (dollars)2 |
Variance = $278,500 (dollars)2.
The variance can be calculated as:
The standard deviation of the initial of $1,000 for Case 2 is . This means that standard deviation of the initial investment of $1,000 is $528 or 53%. The greater the standard deviation, the higher the risk. Case 1 has a standard deviation of $350 and Case 2 has a standard deviation of $528. Case 1 has a lower risk. We can also see this in .

Figure 5.1 Probability distributions showing expected payoffs for both scenarios.
We can see Case 1 is more clustered around the expected value of $1,050 with a standard deviation of 35%, whereas Case 2 is more spread out around the same expected value of $1,050 with a standard deviation of 53%. Case 2 has higher standard deviation, therefore it carries more risk.
Sometimes we are less concerned with spread over possible outcomes from an investment strategy and we are more concerned with what the worst possible outcome may be. Sometimes, we may just want to assess an outcome with a small likelihood of a very large loss. Examples might include: What is the probability that a bank fails? What is the probability of a large loss in the stock market? What is the probability of defaulting on a variable versus a fixed-rate mortgage?
If we are holding stock in a company, the worst-case scenario might be that the company goes bankrupt and its price crashes. This means we take a large loss (this is typically called a “blowout”). The expected value and standard deviation of the price of the stock do not really tell us the risk we face, in this case. VaR answers the question: how much will I lose if the worst possible scenario occurs?
Sometimes this is the most important question we want answered. The following snippet of Python code calculates the maximum loss that could be expected in the next month from a stock portfolio worth $150,000. Two assumptions are required to make the calculation: 1) the variance (volatility) of the portfolio must be assumed or estimated; and 2) we must assume that no further trading takes place within the month.

This shows with 99% confidence that the largest loss that could be expected within the month from the $150,000 portfolio is just shy of $31,000.
Many people are averse to risk—they do not like risk and will pay to avoid it. Insurance is a good example of this. For most buyers, the expected return of their insurance is negative. This means they will pay insurance premiums over their lifetime and never collect a return. However, many find value in insuring against a worst-case scenario by collecting a negative return (in other words, paying) in exchange for minimizing risk.
A risk-averse investor will always prefer an investment with a certain return to one with the same expected return but any amount of uncertainty. Often, investors require compensation for taking on added risk. The compensation investors require to hold the risky asset is called the risk premium. Typically, the riskier the investment, the higher the risk premium. See .

Figure 5.2 Risk premium and expected returns.
There are two main types of risk: idiosyncratic risk and systemic risk.
Idiosyncratic risk refers to that type of risk only affecting a small number of people, a small number of firms or a small sector of industry. These risks are unique to a small portion of an economy. Another type of idiosyncratic risk is a risk that may be bad for one sector of the economy but good for another. For example, a rise in oil prices could be good for the energy industry, while being bad for the car industry.
On the other hand, systemic—or economy-wide risks—are those types of risks that affect a very large swath or even the whole economy. There is a manifest trade-off between risks and returns. Conventionally, chasing higher returns connotes having to deal with higher levels of risk. Therefore, investors must come up with strategies to mitigate risk and plans for responding to both idiosyncratic and systemic risk.
Some investors take on so much risk that a single big loss can wipe them out. Traders call this “blowing up.” Risk can be reduced through a process called diversification, which is the principle of holding more than one risk at a time. Diversification comes in two varieties: hedging and spreading. One can hedge risks or spread them among many investments. If done correctly, these can help reduce the idiosyncratic risk an investor bears.
Hedging is the strategy of reducing idiosyncratic risk by making two investments with opposing risks. Even if one industry is volatile, the payoffs may yet be kept relatively more stable. For example, let us say we have three strategies for investing $100: A—Invest $100 in Exxon (an energy company); B—Invest $100 in Toyota (an automotive company); C—Invest $50 in each company. We will assume oil prices have an equal chance or rising and falling. When oil prices rise, owners of Exxon receive $120 for every $100, while Toyota prices fall to $96. On the other hand, when oil prices fall, owners of Toyota receive $120 for every $100, while Exxon prices fall to $96.
Table 5.4A payoff table for investing $1,000.
| Probability | Expected payoff | Standard deviation |
| Exxon only | 108 | |
| Toyota only | 108 | |
| Half-and-half | 108 |
shows the expected value and standard deviation for each of the three investment strategies. The expected payoff for Exxon-only investment is the same is the same as Toyota’s only, which is . And the standard deviation for either investment is also the same, that is:
What is the expected payoff for hedging with the half-and-half strategy? Instead of investing $100 in either, hedging in this case involves investing $50 in Exxon and $50 in Toyota. Recall there are two states of the world with equal probability: oil prices go up or oil prices go down.
The expected payoff of $50 investment in Exxon is and the expected payoff of $50 investment in Toyota is also , so the total expected payoff of the hedging strategy is also . However, the standard deviation of the hedging strategy is zero:
Importantly, while hedging has effectively eliminated risk from 12% ($12 standard deviation of a $100 investment in our example) in the real world, correct hedging will likely reduce, but will rarely completely eliminate risk.
Investments do not always predictably move in the opposite direction in response to an economic or financial event. Therefore, correct hedging may not always be possible. An alternative risk management approach is to diversify by spreading risk around multiple investments. While hedging required finding investments with opposite risks, this type of diversification requires finding investments with unrelated payoffs to spread the underlying risk.
Let’s now consider a different investment scheme consisting of three strategies for investing $100: A—Invest $100 in American Express (a credit card company); B—Invest $100 in IBM (a technology company); C—Invest $50 in each company. We will assume prices of American Express and IBM move independently. In this case there are four states of the world with equal probability, 25%, of each occurring. Investing $100 in American Express has a 25% probability of returning $120 in state 1, $120 in state 2, $100 in state 3, and $100 in state 4. Investing $100 in IBM has a 25% chance of returning $120, a 25% chance of returning $100, a 25% chance of returning $120, and a 25% of returning $100 in each of the four respective states of the world.
Applying the formulas for expected value and standard deviation that we have seen before, the following snippet of Python code reveals that investing $100 in either company yields the same expected return and the same standard deviation.

If we invest $100 in American Express, the expected value is $110 and the standard deviation is 10%. Similarly, if we invest $100 in IBM, the expected value is $110 and the standard deviation is 10%. Both strategies yield the same expected return of 10% (the expected payoff of $10 from a $100 investment). However, they both have some level of risk, 10%. What if we diversified and invested $50 in each company? The payoff table for the diversified strategy looks as follows:
Table 5.5Calculating the standard deviation of investing $1,000.
| State of the world | Probability | AMEX | IBM | Total payoff |
| #1 | 0.25 | |||
| #2 | 0.25 | |||
| #3 | 0.25 | |||
| #4 | 0.25 |
shows the expected payoff table of the diversified option of investing $50 in American Express and $50 in IBM. According to the table, the expected total payoff of investing $50 in each company is . In this case, diversifying does not alter the expected payoff. Therefore, investing $100 in American Express, investing $100 in IBM, or investing $50 in each yield the same expected payoff.
However, the standard deviation of the diversification strategy is lower than investing in either company, as can be seen below:
The diversified portfolio has a lower standard deviation, so while diversification does not necessarily affect expected payoffs, in this case it reduces risk (from 10% to 7.07%). This can also be seen in Figure .

Figure 5.3 Diversification chart.
The more independent sources of risk held in a portfolio, the lower the overall risk. As we add more and more independent sources of risk, a much lower standard deviation may be achieved, although expected payoffs may also decrease. Diversification through the spreading of risk is the basis for the insurance business, personal financial planning, and investment banking.
Thinking About It…
Risk is defined as uncertainty about an investment’s future payoff. It must be assessed over a time horizon and compared to a benchmark. Risk-free assets have guaranteed returns with no uncertainty.
Expected value of an investment is calculated by summing all possible outcomes multiplied by their probabilities, and it shows the average outcome. However, it does not indicate risk level, so two investments can have the same expected value but different risk levels.
Risk can be better reflected by measures of scatter around the expected value. The variance measures the spread of possible outcomes around the expected value, but the standard deviation (which is the square root of the variance) is more practical as it uses more understandable units of measurement. Higher standard deviations indicate higher risk, and they can be used to compare risk levels between investments.
Risk can be managed through hedging and diversification. Hedging can mitigate risk by making investments with opposing risks. Diversification can help reduce risk by spreading investments across multiple unrelated assets. Proper diversification may reduce overall portfolio risk without necessarily reducing expected returns. Diversification is the foundation for insurance, financial planning, and investment banking.
There is a trade-off between risk and returns. Higher risk typically demands higher expected returns (risk premium). Risk-averse investors prefer certainty, even with lower returns. There are two types of risk: idiosyncratic, affecting specific investments; and systemic, affecting the entire market.
A severe loss scenario where an investment dramatically decreases in value, such as when a company goes bankrupt. Also known as “blowing up” in trading terminology.
In risk analysis, particularly value at risk (VaR) calculations, the statistical probability that a potential loss will not exceed a specified value over a given time period.
A risk management strategy that involves spreading investments across different assets or asset classes to reduce overall portfolio risk. Effective diversification requires investing in assets whose returns are not perfectly correlated.
The weighted average of all possible outcomes of an investment, where each outcome is weighted by its probability of occurrence. It represents the most likely or average outcome if the investment were repeated many times.
A risk management strategy where an investor takes an offsetting position in a related investment to reduce risk exposure. The goal is to protect against adverse price movements by having investments that move in opposite directions under the same conditions.
Also known as unsystematic risk, it refers to risk that affects only a specific asset, company, or small sector of the economy. This type of risk can be reduced through diversification.
Investment returns that are not correlated with each other, important for effective risk spreading in diversification strategies.
The cost paid to transfer risk to an insurance company, typically resulting in a negative expected return for the buyer in exchange for protection against worst-case scenarios.
The total value of a portfolio used as the base for risk calculations, particularly in value at risk (VaR) analysis.
The actual return or value received from an investment in a particular state of the world.
A structured presentation showing all possible investment outcomes and their associated probabilities, used to calculate expected values and risk measures.
A measure of how much a portfolio’s value fluctuates over time, often expressed as a standard deviation and used in risk calculations.
An investment whose future value is known with certainty and provides a guaranteed return. Government securities are often considered the closest approximation to risk-free assets.
The time period over which investment risk is assessed, with longer periods generally associated with higher risk.
The process of identifying, analyzing, and taking steps to reduce or control risk in an investment portfolio through strategies like diversification, hedging, and spreading.
The additional return an investor expects to receive as compensation for taking on extra risk compared to a risk-free investment. Generally, the riskier the investment, the higher the risk premium investors demand.
A statistical measure of the dispersion of possible investment outcomes around their expected value. It provides a standardized way to quantify risk by showing how much returns typically deviate from the average. Higher standard deviation indicates greater risk.
A possible future scenario or outcome that could occur, each with an associated probability and payoff in risk analysis.
Also called systematic or market risk, it represents risk that affects the entire market or economy simultaneously. This type of risk cannot be eliminated through diversification alone.
The duration of time over which an investment or risk analysis is conducted, crucial for proper risk assessment and management strategies.
A statistical measure that quantifies the maximum potential loss an investment portfolio could face over a specific time period at a given confidence level. It answers the question “How much could I lose in a worst-case scenario?”
A statistical measure that quantifies the spread of possible outcomes around the expected value by averaging the squared deviations from the mean. It is the square of standard deviation and serves as a fundamental measure of investment risk.
A statistical measure of the dispersion of returns for a given investment, often used interchangeably with standard deviation in financial contexts.