MTH 451/551
Lecture 2
1 Orders of Convergence
[Textbook reference: §1.2]
In the previous lecture, we saw two types of numerical methods. We saw the midpoint method, a type of quadrature that approximates integrals using a function evaluation at the midpoint of the interval, and finite difference quotients, which approximate derivatives at a point using multiple function evaluations. In both of these examples, we were able to compute the error of the approximation. Recall that the (signed) errors of these approximations are
| (Midpoint Rule) | |||||
| (Forward Difference) | |||||
| (Backward Difference) | |||||
| (Centered Difference) |
where , , and for the midpoint rule, and and for the difference quotients.
In each of these four cases, the error scales in terms of , which we can think of as a parameter. As , then the errors also go to zero, meaning that the approximations become more and more accurate. We mentioned before that the centered difference approximation is more accurate (at least for small ) than either the forward or backward difference approximations. This is because if we halve , then the centered error is (roughly) divided by four, but the forward and backward errors are divided by two. So the centered error shrinks much faster.
We want to study this type of convergence behavior. Consider some quantity that we wish to approximate (for example, the integral or derivative of ), and, for given , an approximation . If the approximation converges, then as . Sometimes we may have approximations such that as , where is usually an integer parameter; we will see examples later (e.g. Gaussian quadrature). We want to quantify and classify how fast converges to as .
Definition 1 (Order of convergence).
Let be a family of approximations of , with as . Given , not necessarily an integer, we say that the convergence is of order when there are constants such that
| (1) |
If is a sequence of approximations to , with as , then we say that the convergence is of order if there exist constants such that
| (2) |
It will be very useful when discussing orders of convergence to use big- notation. This notation is encountered frequently in computer science, when discussing the concept of computational complexity. In our context, the idea is the same, but the setting is slightly different. We write as to express that there exist constants such that for all . When it is clear from context, we will omit the phrase “as ”. So,
Note that if then for all . We prefer to write for the largest for which it holds.
For a sequence , we write as to mean that there exist constants such that for all . So, in the case of as at order , we write .
Note that functions or other than or are commonly encountered. For example, we might see or .
1.1 Estimating Orders of Convergence
As an example, let , and let . We can compute that . How well do the forward, backward, and centered difference quotients approximate ? Let , , and . We know from the expressions for the errors that
Can we verify these results numerically? We compute the difference quotients for , (i.e. successive halvings of ). Let denote the error associated with one of the approximations. Then, for (forward, backward) or (centered). If we assume that for some constant , then
So, for the forward and backward difference quotients, we would expect , and for the centered difference, we would expect .
Slightly more generally, again under the assumption that , if we have two values and , then
If we just have the value of the right-hand side of the above expression, we can recover the value of by computing
We usually use interchangeably with (natural logarithm, base ), but the above expression is independent of the base of logarithm chosen. In the case described earlier, since we considered successive halvings of , we could estimate the order of convergence by computing . We compute these values numerically, and show the results in the table below.
| Forward | Backward | Centered | |||||
|---|---|---|---|---|---|---|---|
| Error | Rate | Error | Rate | Error | Rate | ||
| — | — | — | |||||
Note that if , then . Let , , and . Then, the above tells us that , in other words, there is a linear relationship between the logarithm of the error and the logarithm of , and the slope is the order of convergence. This means that if we plot the error as a function of on a log-log plot, then we expect to observe straight lines with slopes of 1 (forward and backward difference) and 2 (centered difference). This is shown in the plot below.
1.2 Richardson Extrapolation
Richardson extrapolation is a common approach for accelerating convergence when some theoretical knowledge of the structure of the error is known. For example, might be expressed in terms of an expansion in powers of . We motivate the discussion of Richardson extrapolation by revisiting the centered difference approximation of . For simplicity, we assume that is analytic near and that is sufficiently small, so we have
| (3) |
Replacing by in this formula, we obtain
| (4) |
Both and are approximations of , but a simple combination of the two eliminates the term in the error expansions, yielding an approximation. Subtracting from (i.e. subtracting times (4) from (3)) cancels the term, and we obtain
and so
We have taken a linear combination of two instances of the same approximation method for to obtain a new approximation of higher order. This process is called Richardson extrapolation.
Now we will do the same process, but with slightly different notation in order to illustrate the general principle. Write , so
where . We want to cancel the term in the error, which for is and for is , and so we take times minus , and then normalize, giving
where and . Note that the approximation has higher order of convergence, since the error term starts at . The error scales as , so the convergence is fourth order. The important thing to notice here is that this process can now be repeated to cancel the leading order term in the error for . Richardson extrapolation can be applied to where now we need to cancel the term in the remainder. By the same procedure, we need to take and normalize by in the denominator, giving
where and . Now, the error term starts at , so the error for scales as and the convergence is sixth order. This process could be repeated again, canceling the term in the error associated with , and so on. Notice that we did not need to know the coefficients of the error term. We just needed to know the general form of the error. This procedure is formalized in the following theorem.
Theorem 1 (Richardson extrapolation, even powers).
Suppose that, for all sufficiently small, the function approximates the number with error expansion
| (5) |
for some coefficients . Taking , we define the th Richardson extrapolation, , recursively:
| (6) |
If we set for , then
| (7) |
where for .
Proof.
Example 1 (Richardson extrapolation table).
In the context of Theorem 1, we fix (sufficiently small) and define the numbers for . Fixing , we may fill out a Richardson extrapolation table,
It holds that
| (8) |
so only the initial column needs to be computed using ; all other columns may be computed by the given recursion. For each fixed , the entries are all approximations of .
Let , where , , and . We fill out the first column of the table with and use the recursion (8) to complete the Richardson table. The absolute errors , , are given below.
The ratios of successive errors in the initial (zeroth) column quickly approach , as expected, which corresponds to quadratic convergence, . In the next column, the ratios of successive errors are getting closer to the expected value of that corresponds to quartic convergence, . More generally, we have , so we expect as for each fixed .
1.3 Sequential Orders of Convergence
[Textbook reference: §1.3]
Along with the definitions of orders of convergence considered previously, there are a set of different concepts that are also used to classify and quantify how quickly a sequence approaches its limit . We will postpone most of this discussion until after seeing some examples of root-finding methods. However, we will just mention presently that terms such as “linear convergence”, “quadratic convergence,” and so on, have meanings that are different from the “order of convergence” defined previously, even though some authors use terms like linear and quadratic convergence to refer to the previous definition. This can be confusing since similar or identical terms can be used to refer to very different things. We will try to avoid this confusion by being consistent with our terminology, but you may sometimes need to pay attention to the context to understand and correctly interpret which concept is being referred to.