Quantitative Methods · Methods
9Stochastic Control
After a 20% fall in equities a pension fund’s equity weight has drifted from 60% to 52%. Its rulebook says: rebalance to 60% at the quarter-end, which means buying equities after a fall. Half the investment committee wants to wait for the market to settle. The question is one of stochastic control: choose, at each moment and with what is known then, how much to hold, so as to maximise the expected utility of what the fund ends with. For an investor with constant relative risk aversion in a market of lognormal prices, Merton solved it in 1969: hold a constant fraction, whatever has happened, and rebalance to it. With this fund’s assumptions the fraction is 51%, holding 60% instead costs 3.6 basis points of certainty-equivalent return a year, and the committee’s debate is about a much smaller number than the uncertainty in its own expected return. This chapter develops dynamic programming in discrete time, the Hamilton–Jacobi–Bellman equation in continuous time, the verification theorem that turns a candidate into a proof, Merton’s problem, and a sketch of the viscosity solutions that handle value functions with kinks.
9.1 Dynamic programming
Definition 9.1 (Stochastic control problem, admissible control, value function)
A stochastic control problem chooses a process that influences the dynamics of a state , here or its discrete-time analogue, to maximise . An admissible control is an adapted process with values in a set for which the state equation has a solution and is well defined. The value function is over admissible controls.
Theorem 9.2 (Dynamic programming principle)
For every , .
Partial proof. In discrete time with finitely many states and controls: any strategy from earns its rewards up to and then at most , which gives ; following an optimal (or -optimal) continuation from each state at after any first part gives . In continuous time the argument needs measurable selection of -optimal continuations (Fleming and Soner, 2006). ∎
Definition 9.3 (Dynamic programming, feedback control)
Dynamic programming computes the value function backward in time with the one-step principle, , from . A feedback control is a control of the form , a function of the current state; dynamic programming produces one.
The principle reduces a choice over whole strategies to a sequence of one-step choices, at the cost of computing on the whole state space: the curse of dimensionality. The running project’s solver (the tutorial’s first listing) does it on a grid, with the expectation computed by quadrature over the next state and linear interpolation between grid points.
9.2 The Hamilton–Jacobi–Bellman equation
Definition 9.4 (Hamilton–Jacobi–Bellman equation)
For the controlled diffusion of Definition 9.1, the Hamilton–Jacobi–Bellman equation is
It is the dynamic programming principle over an interval with : expand by Itô’s formula, take expectations, divide by . For each fixed control it is the backward equation of chapter 4 with the generator of the controlled state; the supremum makes it nonlinear. The heuristic assumes is smooth, which is what the verification theorem replaces by an assumption about a candidate.
9.3 Verification
Theorem 9.5 (Verification)
Let with polynomial growth solve the HJB equation, and suppose the supremum is attained at such that the state equation with the feedback has a solution. Then and is optimal, provided the stochastic integrals below are true martingales.
Proof. For any admissible , Itô’s formula gives . The HJB equation says , so, taking expectations, : . With the inequality is an equality, so . ∎
9.4 Merton’s problem
Definition 9.6 (Utility function, relative risk aversion, CRRA utility)
A utility function is increasing and concave; the investor maximises . Its relative risk aversion is . CRRA utility has constant relative risk aversion : , and for .
Definition 9.7 (Certainty equivalent)
The certainty equivalent of a random wealth is the sure amount with ; over a horizon , the certainty-equivalent return is .
A fund holds a fraction of its wealth in a stock with and the rest in the money-market account at rate , so .
Theorem 9.8 (Merton)
For CRRA utility of terminal wealth, the optimal fraction is constant,
and the value function is , where is the certainty-equivalent return of the optimal strategy.
Proof. Try in the HJB equation
Since and , the bracket is , maximised at with value . Hence , . The candidate is smooth and the wealth equation with constant is a geometric Brownian motion, so Theorem 9.5 applies. ∎
Definition 9.9 (Merton fraction)
The Merton fraction is .
The optimal fraction does not depend on wealth or on time: after a fall the fund should buy back to it. Holding a constant earns the certainty-equivalent return (Figure 9.1), so a mistaken fraction costs a year. With , , and : , a certainty-equivalent return of 3.29%, and 60% costs 3.6 basis points a year. With , log utility, is the Kelly fraction of One Quant Book 2, chapter 29, here 154%: the growth-optimal investor borrows.
The flatness has a flip side. The fraction is proportional to the expected excess return, which is known far less well: estimated from ten years of returns with 18% volatility, has a standard error of , and every point of moves by ten points (Figure 9.2). Merton’s rule is exact and its input is uncertain; chapter 14 shrinks the input.
9.5 A sketch of viscosity solutions
Value functions are often not smooth: a control that switches from one extreme to another at a boundary puts a kink in , and the HJB equation cannot hold there in the classical sense. The right notion replaces derivatives by smooth test functions touching .
Definition 9.10 (Viscosity solution)
A continuous is a viscosity solution of ( nonincreasing in its last argument) if at every point where a smooth touches from above ( has a local maximum), , and at every point where it touches from below, .
For the problem of leaving as fast as possible at unit speed, and everywhere but at zero, where no derivative exists; it is the viscosity solution, and it is the limit as of the smooth solutions of (Figure 9.3), which is where the name comes from. Crandall and Lions (1983) proved existence and uniqueness for first-order equations; the theory guarantees that dynamic programming on a grid converges to the right value function even when it has kinks, as in the stopping problems of chapter 10.
9.6 Tutorial: Merton by backward induction
Goal. Solve a discrete-time version of the fund’s problem by dynamic programming, check it against Merton’s fraction, and measure what rebalancing buys. End state: Figures 9.1 and 9.2 and the numbers of the weekend problem.
The solver: backward induction on a grid, with an optional stopping reward for chapter 10.
def backward_induction(grid, controls, n_steps: int, reward, transition, terminal, stop=None) -> dict: grid = np.asarray(grid, dtype=float) controls = np.asarray(controls, dtype=float) values = np.empty((n_steps + 1, grid.size)) policy = np.empty((n_steps, grid.size)) stopped = np.zeros((n_steps, grid.size), dtype=bool) if stop is not None else None values[n_steps] = terminal(grid) X, U = np.meshgrid(grid, controls, indexing="ij") for t in range(n_steps - 1, -1, -1): nxt, w = transition(t, X, U) cont = np.interp(nxt, grid, values[t + 1]) @ w # (states, controls) total = reward(t, X, U) + cont best = np.argmax(total, axis=1) values[t] = total[np.arange(grid.size), best] policy[t] = controls[best] if stop is not None: s = stop(t, grid) stopped[t] = s >= values[t] values[t] = np.maximum(values[t], s) return {"values": values, "policy": policy, "stop": stopped}Listing 9.1. Backward induction with quadrature and interpolation. code/firm/dpsolve/firm_dpsolve.py The problem: ten annual rebalancing dates, log wealth on a grid of 241 points, fractions from 0 to 1.5 by steps of 0.005, lognormal annual returns by 20-point Gauss–Hermite quadrature.
def dp_merton(years: int = 10, n_nodes: int = 20, gamma=GAMMA) -> dict: """Annual rebalancing, CRRA utility of wealth after `years`: backward induction over a log-wealth grid with the fraction in the stock chosen on a grid of 0 to 1.5 by steps of 0.005.""" grid = np.linspace(math.log(0.05), math.log(50.0), 241) controls = np.linspace(0.0, 1.5, 301) z, wz = gauss_hermite_normal(n_nodes) gross = np.exp(MU - 0.5 * SIGMA**2 + SIGMA * z) # one year's stock return factor def transition(t, x, u): port = 1 + R + u[..., None] * (gross - 1 - R) # (states, controls, nodes) return x[..., None] + np.log(np.maximum(port, 1e-12)), wz res = backward_induction(grid, controls, years, lambda t, x, u: 0.0 * x, transition, lambda x: crra(np.exp(x), gamma)) inner = (grid > math.log(0.3)) & (grid < math.log(10.0)) return {"policy0": res["policy"][0], "grid": grid, "inner": inner, "pi_dp": float(np.median(res["policy"][0][inner])), "pi_dp_spread": float(np.ptp(res["policy"][0][inner])), "pi_dp_last": float(np.median(res["policy"][-1][inner]))}Listing 9.2. Merton’s problem as a grid dynamic programme. code/methods/09-stochastic-control/python/qm_merton.py - Run
problem(): the optimal fraction is 0.52 at every wealth level and date, the grid point nearest the one-period optimum 0.517 (continuous rebalancing gives 0.514); thenfig_merton.py.
What to change next. Add consumption at a constant fraction of wealth and check Merton’s consumption rule; replace CRRA by a utility with a floor (a liability the fund must cover) and watch the fraction start to depend on wealth.
9.7 Build: the dynamic-programming solver
Purpose. Solve the firm’s small control problems (allocation, inventory, execution schedules) and, with a stopping reward, its stopping problems (chapter 10), by backward induction on a grid.
Interface. backward_induction(grid, controls, n_steps, reward, transition, terminal, stop=None) returning values, policy and the stopping region; gauss_hermite_normal(n).
Rules. One-dimensional state grid, finite control set, expectations by user-supplied nodes and weights, linear interpolation, flat extrapolation; vectorised over states and controls.
Acceptance tests. code/firm/dpsolve/tests/: Gauss–Hermite integrates polynomials exactly; the Merton problem returns a wealth-independent fraction equal to the one-period optimum; a stopping problem with a known solution (the American-style exit of chapter 10) is reproduced.
Stretch. Two-dimensional states with bilinear interpolation; policy iteration for infinite horizons.
Sources and further reading
- R. C. Merton, “Lifetime portfolio selection under uncertainty: the continuous-time case”, Review of Economics and Statistics 51, 1969.
- P. A. Samuelson, “Lifetime portfolio selection by dynamic stochastic programming”, Review of Economics and Statistics 51, 1969.
- R. Bellman, Dynamic Programming, Princeton University Press, 1957.
- M. G. Crandall and P.-L. Lions, “Viscosity solutions of Hamilton–Jacobi equations”, Transactions of the AMS 277, 1983.
- W. H. Fleming and H. M. Soner, Controlled Markov Processes and Viscosity Solutions, Springer, 2nd ed., 2006.
9.8 Exercises
Exercise 9.1 ★
An asset has an expected excess return of 4% and a volatility of 20%. What fraction does an investor with hold?
Solution
Solution of Exercise 9.1.
.
Exercise 9.2 ★
With the fund’s assumptions, what is the certainty-equivalent return of holding 100% in equities?
Solution
Solution of Exercise 9.2.
a year, against 3.29% at the optimum.
Exercise 9.3 ★
Compute the relative risk aversion of and of .
Solution
Solution of Exercise 9.3.
For : . For : .
Exercise 9.4 ★★
Show that holding a constant instead of costs of certainty-equivalent return a year, and evaluate it for 60% against 51.4%.
Solution
Solution of Exercise 9.4.
The certainty-equivalent return is a parabola with maximum at and second derivative , so bp.
Exercise 9.5 ★★
What fraction does a log-utility investor hold with the fund’s assumptions? Why would a fund not do so?
Solution
Solution of Exercise 9.5.
, borrowing 54% of wealth. The fund’s risk aversion is higher than log utility, the expected return is uncertain (overbetting costs more than underbetting, One Quant Book 2, chapter 29), and leverage brings margin and drawdown constraints the model ignores.
Exercise 9.6 ★★
Verify that solves the HJB equation, with , and compute for the fund.
Solution
Solution of Exercise 9.6.
From the proof of Theorem 9.8: and the supremum is , so the equation holds. For the fund, .
Exercise 9.7 ★★★
Coding. With backward_induction, solve the fund’s problem with and annual rebalancing. Is the fraction still independent of wealth, and how does it compare with ?
Solution
Solution of Exercise 9.7.
Yes: 0.315 at every wealth level, against and a one-period optimum of 0.309; the small gap comes from interpolating the value function on the grid.
Exercise 9.8 ★★★
Find the flaw. “We estimated the equity premium at 5% from ten years of data, so our optimal equity weight is 51%, and we will hold exactly that.”
Solution
Solution of Exercise 9.8.
Ten years of returns with 18% volatility estimate the premium with a standard error of , which moves the Merton fraction by points: the data cannot distinguish 0% from 100%. Shrink the estimate toward a prior (chapter 14) and size for the uncertainty, not for the point estimate.
9.9 Problem: The Rebalancing Committee
Problem 9.1
Weekend problem — how much the committee’s decision is worth
A fund with CRRA utility, , expects equities to return with volatility ; cash earns . Its rulebook holds 60% in equities, rebalanced. After a fall its weight is 52%.
Part I — Merton.
- What is the Merton fraction?
- What is the certainty-equivalent return of the optimal strategy?
- What does holding 60% cost a year?
- And holding 52%?
- What trade does Merton recommend after the fall, and what does the rulebook recommend?
Part II — Sensitivity.
- What are the fractions for and ?
- And for and ?
- What does give, and what does it mean?
- What does the grid dynamic programme with annual rebalancing give, and the one-period optimum?
- Why does the fraction not depend on wealth?
Part III — Rebalancing.
- Over ten simulated years, what certainty-equivalent returns do constant 60%, Merton, and buy-and-hold from 60% achieve?
- Where does the buy-and-hold weight end after ten years (10th percentile, median, 90th)?
- What does never rebalancing cost, against rebalancing to 60%?
- With what standard error is known from ten years of returns, and what band does that put on ?
- Which matters more for the fund: the committee’s choice between 52% and 60%, or the uncertainty in ?
Part IV — Judgement.
- What would transaction costs change?
- What would liabilities (a pension promise) change?
- What should the chief investment officer tell the committee?
- State the named result: the Merton fraction and the cost of the rulebook’s 60%.
- In one sentence: what does Merton’s solution say about rebalancing after a fall?
Solution
Solution of Problem 9.1.
1. . 2. . 3. 3.6 bp a year. 4. 0.02 bp: 52% is almost exactly optimal. 5. Merton: sell 0.6 point, practically nothing; the rulebook: buy 8 points back to 60%. 6. 77% and 31%. 7. 41% and 62%. 8. 154%: the Kelly, growth-optimal fraction, with borrowing. 9. 0.52 at every wealth level and date; the one-period optimum is 0.517 (continuous: 0.514). 10. CRRA utility is homothetic: scaling wealth scales utility, so the best fraction is the same at every wealth. 11. 3.28%, 3.31% and 3.24% (theory for Merton: 3.29%). 12. 50%, 68% and 81%. 13. 3.3 bp a year. 14. 5.7%; moves by points either way, from about to . 15. The uncertainty in , by far: the committee is debating a few basis points of certainty equivalent. 16. A band around the target inside which the fund does not trade (chapter 10). 17. The fund would hold a liability-hedging portfolio plus a Merton-like position whose size depends on its funding surplus; the fraction would depend on wealth. 18. Rebalance mechanically to a target; spend the debate on the inputs (expected return, risk aversion) and on costs, not on timing. 19. Named result: the rebalancing committee: the Merton fraction is 51.4%; holding the rulebook’s 60% costs 3.6 bp a year of certainty-equivalent return, never rebalancing about 3.3 bp, and one standard error in the expected return moves the fraction by 59 points. 20. Rebalance back to the constant fraction, since nothing about the optimum changed with the price.
9.10 Interview questions
Interview question 9.1 ★ researcher, trader
State Merton’s optimal fraction and explain each of its three inputs.
Solution
Solution of Interview question 9.1.
: the expected excess return (more reward, more stock), risk aversion (more aversion, less), and variance (more risk, less): a mean-variance trade-off per unit of wealth.
What the interviewer is looking for: the formula and its comparative statics.
Interview question 9.2 ★★ researcher
Derive the HJB equation from the dynamic programming principle.
Solution
Solution of Interview question 9.2.
; expand by Itô, take expectations, subtract , divide by and let : .
What the interviewer is looking for: Itô inside the dynamic programming principle.
Interview question 9.3 ★★ researcher
Why does a CRRA investor in Merton’s model buy equities after they fall?
Solution
Solution of Interview question 9.3.
The optimal fraction is constant; a fall lowers the equity share below it, so restoring it means buying. Nothing in the model’s opportunity set changed with the price: returns are independent and identically distributed.
What the interviewer is looking for: constant fraction plus i.i.d. returns.
Interview question 9.4 ★★ trader, researcher
How does the Kelly criterion relate to Merton’s problem?
Solution
Solution of Interview question 9.4.
Kelly maximises expected log wealth: Merton with , giving . More risk-averse investors hold a fraction of it: fractional Kelly is Merton with .
What the interviewer is looking for: log utility and fractional Kelly.
Interview question 9.5 ★★ developer
Your dynamic programme on a grid returns a policy that jumps between neighbouring grid points. What do you check?
Solution
Solution of Interview question 9.5.
Whether the objective is flat there (then the choice barely matters), the control grid is too coarse, the quadrature too crude, or the interpolation of the value function too rough; refine each and see whether the policy converges.
What the interviewer is looking for: flatness versus numerical error, and a convergence check.
Interview question 9.6 ★★★ researcher
When is the value function of a control problem not smooth, and how does one make sense of the HJB equation then?
Solution
Solution of Interview question 9.6.
When the optimal control switches between extremes, or at a stopping boundary, or with degenerate diffusion: the value function has kinks. It is then the unique viscosity solution: at every point, smooth functions touching it from above or below satisfy the equation’s inequalities, which is also the notion under which grid schemes converge.
What the interviewer is looking for: test functions touching from above and below.