MTH 451/551
Lecture 4
1 Solving Equations and Root Finding
[Textbook reference: §2.1]
We are interested in finding the solution to the system for some function , i.e. the roots of . Depending on the specific form of the function , we may, through some complicated algebraic manipulations, be able to find a closed-form expression for the roots. However, this is a laborious process, and we are not always guaranteed to succeed. For example, as mentioned at the beginning of the course, there is no general closed-form expression (written in terms of radicals) for the roots of polynomials of degree 5 or higher. Therefore, it is of great utility to develop algorithms that approximate the roots, using only evaluations of (and perhaps its derivatives), without symbolic or algebraic manipulations.
1.1 Root Bracketing Methods
Two of the oldest and most basic methods for root finding are the bisection and regula falsi methods, which are both root bracketing methods. This means that if there is a root of in the interval (we say that and bracket the root, since they lie on either side of it), then we generate a nested sequence of intervals , each of which contain the root . Within each interval, we will obtain an approximate root that converges to as .
1.1.1 The Bisection Method
[Note: we differ slightly from the textbook here by using closed intervals and non-strict inequalities.]
We consider a continuous function , such that . This means that either and have opposite signs, or at least one of them is zero. By the intermediate value theorem, this implies that there exists a root of in the interval . The bisection method generates a sequence that converges to a root .
-
1.
Initialize , .
-
2.
Iterate .
To turn this into a practical algorithm, we need to set some termination criteria. The sequence , and we need to decide when to stop iterating, and accept the current value of as a good enough approximation to (very rarely we may find such that exactly, but this does not happen usually). Typically, we set both a tolerance and maximum number of iterations. In the version of the algorithm presented below, we provide two tolerances and a maximum number of iterations. We terminate if either or if . Even though we don’t know the exact value of (if we did, we wouldn’t need an algorithm to approximate it), we can bound since is guaranteed to lie in every interval we generate.
Pseudocode version of the bisection algorithm: return an approximation of a root of , given that .
The following theorem guarantees that the bisection method always converges to a root, and gives an a priori bound on the error at each iteration.
Theorem 1 (Convergence of the bisection method; Textbook Theorem 2.2).
Suppose that with . If is generated by the bisection method, then , where . Furthermore, we have the error bound
Proof.
First, we show by induction that for all . This holds for by assumption. Suppose it holds for some . If , then , and so . Otherwise, , so and are nonzero and have the same sign. Since , it follows that , and so .
By construction, , and the intervals are nested, so
for all . The sequence is nondecreasing and bounded above by , and the sequence is nonincreasing and bounded below by . By the Monotone Convergence Theorem, there exist such that and . Furthermore, since each step halves the length of the interval, , we see that
for all . Taking the limit as , we find , and so . We take to be this common value.
Since is continuous at , we have . Since for all , the limit must also satisfy . But clearly , so we must have .
Finally, since is the midpoint of and , the distance from to is at most half the length of the interval,
The right-hand side tends to zero as , and so . ∎
1.1.2 Regula Falsi
The regula falsi method (latin for false position) is a simple variation on the bisection method. Like bisection, it is a root bracketing method. The “false positions” in this case are the two endpoints, which are generally not roots themselves, but are used to help find the roots. Bisection proceeded by always choosing the midpoint of the interval, leading to a halving of the error bound at every step. This had no dependence on the function whatsoever. Regula falsi, on the other hand, chooses to be the zero of the secant line connecting and . This uses more information about the function , and can often converge much faster than bisection, but lacks the same guaranteed error reduction rate.
The setup is the same as in bisection, except here we require that strictly in order for the secant line to have nonzero slope. The method is summarized as follows:
-
1.
Initialize , .
-
2.
Iterate .
The formula for is the zero of the secant line through and . It can equivalently be written as
The termination criteria are the same as for bisection, with one important difference. Since and both lie in , we can still bound the error by . However, unlike in bisection, the length of the interval does not necessarily go to zero. Often, one of the endpoints will stay fixed for all iterations. In that case, this error bound will never drop below , and the iteration will be terminated by (or ) instead.
Pseudocode version of the regula falsi algorithm: return an approximation of a root of , given that .
The following theorem guarantees that the regula falsi method always converges to a root. In contrast to the bisection method, there is no a priori bound on the error.
Theorem 2 (Convergence of the regula falsi method; Textbook Theorem 2.7).
Suppose that with . If is generated by the regula falsi method, then , where .
Proof.
We only consider the non-trivial case, where for all . We lose no generality by assuming that and , which guarantees that and for all . In particular, is well-defined, and since the secant line changes sign between and , we have .
As in the proof for bisection, is nondecreasing and is nonincreasing, so and for some . By continuity, . If , then we take to be this common value, and and by the same argument as for bisection.
Suppose instead that . Since is either or , no lies in the interval . If , then by continuity,
If , then , and we take . Similarly, if , then , and we take . If both and , then lies strictly between and , which is impossible, since no lies in . The only remaining possibility is with . This is clearly impossible if has only one root in ; the general case is more involved (see textbook §2.1, Exercise 13). ∎