MTH 451/551
Lecture 3
1 Floating Point Numbers and Arithmetic
[Textbook reference: §1.4]
As discussed in the first lecture, the point of numerical analysis is to use computer algorithms to solve mathematics problems. Therefore, it is supremely important to understand how numbers are represented by the computer.
1.1 Prelude: Integer Representations
For integers, there is no real ambiguity. Integers are represented by the computer using their representation in binary. Typically, integers are stored using a fixed number of bits (for example, 32 or 64), which allows us to represent all integers in a finite range. A bit is simply a placeholder that can take the value 0 or 1. Representing integers beyond that range (i.e. with larger magnitude) simply requires using more bits. As an example, if we are encoding non-negative integers, we can represent the numbers 0, 1, 2, …, 7 using three bits. This is a total of representable numbers, corresponding to ways to choose from three times.
| Number | Binary Representation |
| 0 | 000 |
| 1 | 001 |
| 2 | 010 |
| 3 | 011 |
| 4 | 100 |
| 5 | 101 |
| 6 | 110 |
| 7 | 111 |
The number associated with the binary representation is . If we wished to represent signed numbers (positive or negative), we would use one bit of information to encode the sign (there are various ways of doing this). The upshot of this discussion is that it is relatively straightforward and uncontroversial how to represent integers on the computer.
1.2 Floating Point Numbers
As soon as we want to represent real numbers (or even just rational numbers), the situation becomes more complicated. As opposed to the integer case, where a given number of bits can represent all possible integers within a range, the same is fundamentally impossible for real numbers, because every interval contains infinitely many of them, and so they cannot all be representable with finite information. This means that, even when representing numbers within a given range, we need to consider issues of precision and rounding. Floating point numbers are the most commonly (almost universally) used method to represent these kinds of numbers on the computer.
Before we define floating point numbers, it is useful to contrast them with the conceptually simpler but generally less useful fixed point numbers. We can think of fixed point numbers as “fixing the decimal point” in an integer number, allowing us to represent fractional numbers. Essentially, this is the same as fixing a denominator , and then just representing integers, which get multiplied by . These types of numbers are often used in financial calculations, where quantities in dollars are represented using a denominator of , allowing the exact representation of all dollars-and-cents quantities within a range. However, outside of this and some other niche applications, fixed point numbers are relatively rarely used, in favor of the much more flexible floating point numbers. The key advantage of floating point numbers is that, essentially, the fractional multiplier is also encoded as part of the number. This way, rather than “fixing” the position of the decimal point, it is allowed to “float”.
Floating point numbers divide up the bits to represent three parts of the number: the sign, the mantissa, and the exponent. These terms correspond to representing a number as
where encodes the sign (either 0 or 1), is the mantissa, and is the exponent; here, is called the base of the representation. This is essentially the same as “scientific notation” for numbers, which you may be familiar with. So, to represent a floating point number, we need to encode , , and . On the computer, it is natural to use binary (base two) representation, and we take .
The sign is represented using a single bit. The number of bits used to represent the mantissa and exponent depends on the exact choice of representation. Luckily, these representations have been standardized in a standard called IEEE 754 that is basically universal among all computers. This standard defines how to represent floating point numbers with 32, 64, or 128 bits.
| Total Bits | Sign | Mantissa | Exponent |
|---|---|---|---|
| 32 | 1 | 23 | 8 |
| 64 | 1 | 52 | 11 |
| 128 | 1 | 112 | 15 |
32-bit floating point numbers are called single precision. Likewise, 64-bit floating point numbers are called double precision, and 128-bit floating point numbers are called quadruple precision. In numerical analysis, double precision is the most widely used. In other fields, such as machine learning, single precision (or even half precision, using 16 bits) is more widely used; similarly for computer graphics. In this course, we will focus on 64-bit double-precision floating point numbers.
The exponent is represented by 11 bits, so it can take on a total of possible values; these give the range , along with two reserved values used for special numbers. The 52-bit mantissa is given by
where is the value of the th mantissa bit. The subscript “2” in the expression indicates that the digits after the “point” are interpreted in base 2. The sign bit can be either 0 or 1. The number represented by these 64 bits is
We use the symbol to denote the (finite) set of all such 64-bit numbers, together with the number zero. (Note: these are the normalized floating-point numbers. Some special numbers, such as positive infinity and negative infinity, as well as so-called “subnormal” numbers, can also be represented in this format, but these are beyond the current scope).
The mantissa is always in the interval . We can be more precise. The smallest mantissa is when all its bits are zero, and is exactly equal to 1. The largest mantissa is when all its bits are one, and is equal to
The spacing between successive mantissa values, which is called machine epsilon, is
The spacing between consecutive floating point numbers in is (and likewise for ). For this reason, and as the next theorem will demonstrate, it is often said that double precision has (approximately) 16 (decimal) digits of precision.
Another special value in is called unit roundoff, and is given by
This number features in error analysis of rounding and floating point arithmetic. The smallest and largest positive numbers in are
We say that a real number is in the range of if or if . For such an , we define the rounded value of to be the number nearest to . In case of a tie, we round to the nearest element of with (we say that ties round to even; this is an unbiased rounding method).
The following theorem bounds the relative error between a number in the range of and its rounded value , in terms of the unit roundoff .
Theorem 1.
If is in the range of , then for some . In other words, for non-zero , the relative error in rounding is smaller than ,
Proof.
Without loss of generality, we may assume since if then and we may take . The first step is to write as
where is the mantissa and is the exponent. All nonzero real numbers can be written this way. If , then again and , so we now consider only the case that . As remarked above, the floating point numbers between and are evenly spaced, with spacing (and the same for the floating point numbers between and ). So, lies in some interval , where both endpoints are in . The closer of the two endpoints is (and if there is a tie, apply the round-to-even rule). The distance from to the closest of the endpoints is at most half the interval width, i.e. . Consequently, . Since , it follows that
Let . It can be seen immediately that . The above inequality implies that . ∎
Note that even some simple rational numbers such as are not in . To see this, recall that for , we can compute the limit of the geometric series,
Let for natural number . Then,
This is an infinitely repeating pattern of digits after the point. If , then the digits are not all 1. Since every number in has a finite binary expansion, this implies that . As an example,
The bolded number (which is equal to the sum in parentheses) is the largest number in that is less than . The closest two floating point numbers above and below are
We can see that
and so is twice as close to as is. Therefore,
Note that every number in has an exact decimal representation, but it may take many decimal places to write it out in full. In general, only the first of those digits will be meaningful.
1.3 Floating Point Arithmetic
We have now defined which numbers we can represent on the computer: those in (as well as some “special” numbers mentioned above, which we will not discuss further). We now need to define how to perform arithmetic on these numbers. The four fundamental arithmetic operations are , , , . In this section, we will use to generically denote one of these operations, i.e. to serve as a stand-in for any of them.
Even if , it may be that . This is clear for division. Even though and , we saw earlier that . It is also true for the other three operations. For example, and , but .
Let denote the floating point versions of each of the operations, denoted generically by . These are functions . (Division by zero requires a special definition; we avoid discussing this case). For , we define
This means that the computed value is given by rounding the exact value to the nearest floating point number. This same property is required or suggested for other operations, such as . Function such as etc. are not required to give the rounded version of the exact value. This will be evident in some of the examples we consider.
While some familiar properties of standard arithmetic operations still hold for their floating point equivalents, other properties do not hold. For example, addition and multiplication are commutative in both exact arithmetic and in floating point arithmetic:
However, floating point operations are not generally associative. To illustrate this, consider (in exact arithmetic),
However, in floating point arithmetic, note that
and so
but
Note that is expected to have correct digits, and so the basic arithmetic operations are expected to be computed with a high degree of accuracy, since . However, these small errors can accumulate, sometimes in quite significant ways, leading to calculations that have much larger relative errors. This is often the result of subtracting two distinct but very close numbers.
Example 1.
Let . For large , the exponent is very close to zero, and so is very close to one. Let . We compute for various values of , and show the result in the table below.
| 25 | |||
|---|---|---|---|
| 26 | |||
| 27 | |||
| 51 | |||
| 52 | |||
| 53 | |||
| 54 | |||
What we see from this table is that for ,
This means that the relative error is
For this calculation, there are no correct significant digits.
The previous example may seem somewhat contrived. The next example shows that in practical situations, being aware of potential round-off errors can be extremely important.
Example 2.
Recall the forward difference approximation to the derivative,
Take and . We know that , and so . We compute the difference quotient in floating point arithmetic. We expect that the error will go to zero linearly as . However, for very small values of , the error from subtractive cancellation can be catastrophic. From the table in the previous example, we see that for , . Therefore, for ,
and the relative error is again 100%. We can also see from the previous table that for , the relative error is 100%. In this case, the numerator is and the denominator is , and so the difference quotient gives an approximation of , but the exact answer is 1. The table below shows the finite-difference approximation and relative error for values of .
| Relative Error | ||
|---|---|---|
| 22 | ||
| 23 | ||
| 24 | ||
| 25 | ||
| 26 | ||
| 27 | ||
| 52 | ||
| 53 | ||
| 54 |
The upshot is that, while the finite difference error decreases as , we cannot take too small, otherwise round-off error will destroy the accuracy.
Note that the exact values obtained in the example above (where the relative error is zero) are essentially coincidences because of the particular choice of and . If we choose , then we obtain the following results.
| Relative Error | ||
|---|---|---|
| 16 | ||
| 18 | ||
| 20 | ||
| 22 | ||
| 24 | ||
| 26 | ||
| 28 | ||
| 30 | ||
| 32 | ||
| 34 | ||
| 36 | ||
| 38 | ||
| 40 | ||
| 50 | ||
| 51 | ||
| 52 | ||
| 53 |
This example shows the most common behavior of the error for finite difference quotients: at first, the error decreases with , since the term gets smaller. Then, for sufficiently small , the round-off error becomes, significant, and the error starts to increase until it completely dominates. Notice that the error achieved a minimum of about , which is very far from the 16 decimal digits of accuracy we would hope to obtain. We can help address this issue by using a centered difference approximation, which has error .
| Relative Error | ||
|---|---|---|
| 8 | ||
| 10 | ||
| 12 | ||
| 14 | ||
| 16 | ||
| 18 | ||
| 20 | ||
| 22 | ||
| 24 | ||
| 26 | ||
| 28 | ||
| 30 | ||
| 32 | ||
| 34 | ||
| 50 | ||
| 51 | ||
| 52 | ||
| 53 |
Here, we see that the error achieves a minimum of about , which is 10,000 times smaller than in the case of forward differences.
We can predict roughly what error is achievable by considering both sources of error: the truncation error from the Taylor expansion, and round-off error from the floating point arithmetic. The truncation error scales like . The round-off error in the numerator will scale like , and so after dividing by the denominator (and taking the relative error, with exact value ), we see that the contribution to the relative error from round-off scales like . So, the overall error will scale like .
We can find the that minimizes . To do this, we find the critical point . This gives . This model predicts that the error will be minimized at and the error will be bounded roughly . Indeed, in our table, the minimum is achieved at , and the observed error is less than the bound provided by the model.
Example 3.
For functions which are analytic (complex differentiable around ), and which take real values on the real line, then we can use complex variable methods to get even more accurate approximations to derivatives. Counterintuitively, these methods use complex numbers to approximate results that are always real numbers. If we evaluate the Taylor series of about at the point (where is the imaginary unit), we obtain
and so
This is a second-order approximation requiring only one (complex) evaluation of . Furthermore, there is no subtraction in the numerator, and since the imaginary part of is , the round-off error in computing will scale like . Therefore, we can model the error as . The round-off term does not grow as , allowing us to use extremely small values of and still obtain accurate results.
| Relative Error | ||
|---|---|---|
| 8 | ||
| 10 | ||
| 12 | ||
| 14 | ||
| 16 | ||
| 18 | ||
| 20 | ||
| 22 | ||
| 24 | ||
| 26 | ||
| 28 | ||
| 30 | ||
| 32 | ||
| 34 | ||
| 200 | ||
| 201 | ||
| 202 | ||
| 203 |
From these results, we see that the complex step method achieves as good accuracy as is possible, without needed to balance the competing forces of truncation error and round-off error.
Our final example concerns linear algebra.
Example 4.
For , consider the solution of the following linear system by Gaussian elimination with back-substitution,
which yields
For , both and are very nearly 1. However, if we evaluate these formulas in floating point arithmetic, we obtain (which is in fact , the best possible value), which then gives . The relative error in is therefore , i.e. essentially 100%.
This example illustrates that simple linear algebra problems where all entries of the matrices and vectors can be represented exactly in floating point (and the solution can be represented very accurately in floating point) can still lead to catastrophic round-off errors.
However, through careful design of linear algebra algorithms, these problems can be avoided. Such algorithms are called numerically stable methods. Indeed, if we use a standard Gaussian elimination routine on the computer to solve this problem, we will obtain answers that are extremely accurate. The methods used are specifically designed to avoid these bad round-off errors.