Machine Learning for Markets · Machine learning
19Deep Hedging and Machine Learning in Pricing
A one-month at-the-money call on a stock whose volatility follows Heston’s model is worth 2.25 on a spot of 100, and one point of implied volatility is worth 0.115 of it. Hedged every day to its Black–Scholes delta with a transaction cost of 5 basis points, the hedge pays 0.123 in costs on average: more than a volatility point, for a trader whose edge is often less. A network that learns when not to trade pays 0.099 and carries less risk. This chapter uses machine learning where derivatives desks use it: to hedge under frictions that the textbook ignores (deep hedging), to replace a slow pricer by a fast approximation (surrogates, and the differential learning that makes them accurate with few samples), and to calibrate a model in one pass. Each is measured against the exact answer that Book 5’s engines provide, and each has a place where it breaks.
19.1 Hedging under frictions as a learning problem
Book 5 (chapter 4) measured the discrete hedging error of the Black–Scholes delta and (chapter 26) the hedging band that transaction costs call for. With costs, the best hedge depends on the current holding (trading back to delta is not worth it if delta has moved little), on the risk the trader is willing to carry, and on everything the model leaves out; there is no closed form beyond asymptotic ones. The problem is a stochastic control problem (Book 4, chapter 9) whose objective is a risk measure of the hedged P&L.
Definition 19.1 (Convex risk measure, entropic risk measure, indifference price)
A convex risk measure assigns to a random P&L the amount of cash that makes it acceptable, and is monotone, cash-invariant () and convex (Föllmer and Schied, 2002); expected shortfall (Book 6, chapter 21) is one. The entropic risk measure is , the certainty equivalent of exponential utility with risk aversion . The indifference price of a claim is the premium that leaves the seller exactly as well off, by the risk measure, as not selling, each with its best trading strategy: , where is the gain of trading strategy after costs. When the underlying’s price is a martingale, as in the chapter’s simulation, not trading is best without the claim and the second term is zero.
19.2 Deep hedging
Definition 19.2 (Deep hedging)
Deep hedging represents the hedging strategy by a neural network that maps what is observable at each date (time, prices, current holdings) to the new holdings, and trains it on simulated paths to minimise a convex risk measure of the hedged P&L after all frictions (Buehler, Gonon, Teichmann and Wood, 2019).
The chapter’s network (Listing 19.1) is the no-transaction band network of Imaki and co-authors (2021): from the time to maturity and the log-moneyness it outputs a band, and the new holding is the old one clipped into it, so that not trading is the default and the band’s width can be read off. The P&L of selling the call and hedging daily with the underlying is computed on a batch of paths (Listing 19.2), the entropic risk with is the loss, and 1 000 Adam steps on batches of 2 048 of 20 000 training paths take about fifteen seconds on one core. Every policy is then scored on 20 000 other paths.
ml_hedge.hedging.Figure 19.1 is the first half of the chapter’s named result. Without costs the deep hedge’s indifference price is 2.454 against 2.475 for the Black–Scholes delta: under Heston, spot and volatility move together () and the network learns a hedge ratio that uses it. With costs of 5, 10 and 20 basis points it rises to 2.558, 2.658 and 2.840; the Black–Scholes delta hedge’s to 2.608, 2.743 and 3.018, and the Whalley–Wilmott band’s (Whalley and Wilmott, 1997) to 2.606, 2.712 and 2.910. Table 19.1 shows the other statistics at 10 basis points. Trained on expected shortfall instead, the network lowers the expected shortfall (3.763 against 3.817) and raises the entropic risk (2.671 against 2.658): the risk measure is part of the model, and each network is best by its own.
| policy | entropic risk | expected shortfall 95% | mean cost | standard deviation |
|---|---|---|---|---|
| deep hedge, entropic | 2.658 | 3.817 | 2.442 | 0.643 |
| deep hedge, expected shortfall | 2.671 | 3.763 | 2.457 | 0.657 |
| Whalley–Wilmott band | 2.712 | 4.044 | 2.424 | 0.703 |
| Black–Scholes delta | 2.743 | 4.090 | 2.507 | 0.595 |
| no hedge | 8.767 | 9.992 | 2.235 | 2.945 |
ml_hedge.hedging, shortfall_hedge.The second half of the named result is the band itself (Figure 19.2). At the money half-way to maturity, the learned band is 0.031 wide in delta at 5 basis points and 0.104 at 20; the Whalley–Wilmott band, , is 0.181 and 0.288. The asymptotic formula assumes a trader who watches continuously and trades at the band’s edge; with one decision a day, the position drifts between decisions and a wide band compounds the drift, so the network keeps a narrower band at the money and above it, and a wider one below the money, where the option is nearly worthless to hedge. The formula gives the right idea and the wrong size for this trader; the network finds the size for the trader’s actual schedule and risk measure.
ml_hedge (fig_hedge.py).19.3 Surrogate pricers and differential learning
Definition 19.3 (Surrogate model, differential machine learning)
A surrogate model is a fast approximation of a slow function, here a pricer, fitted to its outputs on sampled inputs. Differential machine learning trains the surrogate on the derivatives of the labels with respect to the inputs as well as on the labels, the network’s own derivatives being computed by automatic differentiation (Book 4, chapter 28); with Monte Carlo labels the derivative labels are pathwise (Huge and Savine, 2020).
A risk system that prices a book under thousands of scenarios cannot afford a Fourier integral or a Monte Carlo run per price. The chapter’s surrogate learns the Heston price of a three-month call as a function of the spot (70 to 130) and the initial variance (0.01 to 0.09) from the cheapest labels there are: one simulated path per sample, whose payoff is an unbiased but very noisy price. The differential label is the pathwise derivative of the payoff with respect to the spot, , free once the path exists. Table 19.2 scores both trainings against the Fourier price on 400 test points (Listing 19.3 is the training loop).
| price RMSE | delta RMSE | |||
| training samples | standard | differential | standard | differential |
| 256 | 0.927 | 0.679 | 0.066 | 0.057 |
| 1 024 | 0.436 | 0.218 | 0.059 | 0.021 |
| 4 096 | 0.297 | 0.279 | 0.042 | 0.020 |
ml_hedge.surrogates.With 1 024 samples the derivative labels halve the price error and cut the delta error by nearly two-thirds; with 4 096 the prices are as good either way but only the differential network’s deltas are. A surrogate’s delta is what the hedger uses, and a network fitted to prices alone has no reason to get its slope right.
19.4 Calibration networks
Definition 19.4 (Deep calibration)
Deep calibration replaces the numerical optimisation of a model’s parameters against market quotes (Book 4, chapter 24) by a network trained on model-generated data: either the inverse map from quotes to parameters (Hernandez, 2016) or a fast forward map from parameters to quotes inside a standard optimiser (Horvath, Muguruza and Tomas, 2021).
The chapter’s network learns the inverse map from fifteen normalised call prices (five strikes, three maturities) to the five Heston parameters, on 4 000 parameter sets drawn uniformly in a box. On 200 test sets, its errors as a share of each parameter’s range are 1.7% for the initial variance, 18.7% for the mean-reversion speed, 9.7% for the long-run variance, 9.0% for the volatility of variance and 9.8% for the correlation; the surfaces re-priced with its answers miss the true ones by 13.2 basis points of spot on average. Speed and long-run variance are poorly identified by three maturities, and the network is as uncertain about them as the data are. On five test surfaces its answer misses the implied volatilities by 0.46 points on average, and Book 5’s Levenberg–Marquardt calibrator, started from it or from a fixed guess, fits them exactly. The network is a fast first guess and a sanity check, not a replacement, on clean data from the model it was trained on; on market data, which no model fits exactly, it has never seen the residuals it will meet.
19.5 Where learned pricing breaks
Three failures recur. Extrapolation: a surrogate or a calibration network outside its training box returns confident numbers with no warning, exactly where a stressed market takes the book; train on boxes wider than any scenario and check inputs against them. Arbitrage: a surrogate’s prices need not be monotone or convex in the strike, and its Greeks can change sign where they should not; constrain the architecture or check the outputs. Model risk: a deep hedge is optimal for the simulator it was trained on, like chapter 17’s agent; train it on several models or randomised parameters, and compare it with the asymptotic band on paths it was not trained on.
Method 19.5 (Machine learning on a derivatives desk)
- Keep the exact pricer as the reference; every surrogate is validated against it on a grid wider than the book’s scenarios, prices and Greeks both.
- Train surrogates with differential labels when derivatives are cheap (pathwise or by automatic differentiation).
- Use calibration networks as starting points and checks; the optimiser has the last word.
- Train deep hedges on the desk’s own frictions and risk measure, and compare with the closed-form band out of sample and across models.
19.6 Tutorial: hedging with frictions
Goal. Deep-hedge the call at four cost levels and read off the indifference prices and bands; train the surrogate with and without differential labels; fit the calibration network. End state: Figures 19.1 and 19.2, Tables 19.1 and 19.2.
The no-transaction band network.
class BandNet(nn.Module): """Inputs (time to maturity / T, log-moneyness); outputs a band [lower, upper] in delta units. The new holding is the previous one clipped into the band.""" def __init__(self, hidden=(32, 32)): super().__init__() layers, d = [], 2 for h in hidden: layers += [nn.Linear(d, h), nn.ReLU()] d = h self.net = nn.Sequential(*layers, nn.Linear(d, 2)) def forward(self, tau_frac, logm, prev): o = self.net(torch.stack([tau_frac, logm], -1)) centre, half = torch.sigmoid(o[..., 0]), nn.functional.softplus(o[..., 1] - 3.0) return torch.minimum(torch.maximum(prev, centre - half), centre + half), centre - half, centre + halfListing 19.1. A network that outputs a band; the holding is clipped into it. code/firm/deephedge/firm_deephedge.py The hedged P&L after costs.
def hedge_pnl(policy, S, K, T, cost): """Selling one call and hedging at each of the n dates before maturity: the terminal P&L (without the premium) = sum_k delta_k (S_{k+1} - S_k) - cost * sum_k |delta_k - delta_{k-1}| S_k - cost * |delta_{n-1}| S_n - (S_n - K)+. policy(tau_frac, logm, prev) -> (holding, lower, upper).""" S = torch.as_tensor(S, dtype=torch.float32) n = S.shape[1] - 1 prev = torch.zeros(S.shape[0]) pnl = torch.zeros(S.shape[0]) for k in range(n): tau_frac = torch.full_like(prev, (n - k) / n) h, _, _ = policy(tau_frac, torch.log(S[:, k] / K), prev) pnl = pnl + h * (S[:, k + 1] - S[:, k]) - cost * torch.abs(h - prev) * S[:, k] prev = h pnl = pnl - cost * torch.abs(prev) * S[:, n] - torch.clamp(S[:, n] - K, min=0.0) return pnlListing 19.2. Selling a call and hedging it on a batch of paths. code/firm/deephedge/firm_deephedge.py Differential training.
def fit_surrogate(X, y, dy=None, differential=False, seed=0, epochs=400, lr=5e-3, weight=1.0, hidden=(64, 64)): """Fit price labels y (and, if differential, derivative labels dy with respect to the first input) by mean squared error, the derivative of the network taken by automatic differentiation (Huge and Savine).""" torch.set_num_threads(1) torch.use_deterministic_algorithms(True) torch.manual_seed(seed) net = SurrogateNet(X.shape[1], hidden) Xt = torch.as_tensor(X, dtype=torch.float32) yt = torch.as_tensor(y, dtype=torch.float32) dyt = None if dy is None else torch.as_tensor(dy, dtype=torch.float32) opt = torch.optim.Adam(net.parameters(), lr=lr) for _ in range(epochs): x = Xt.clone().requires_grad_(differential) p = net(x) loss = ((p - yt) ** 2).mean() if differential: dp = torch.autograd.grad(p.sum(), x, create_graph=True)[0][:, 0] loss = loss + weight * ((dp - dyt) ** 2).mean() opt.zero_grad() loss.backward() opt.step() return netListing 19.3. A surrogate trained on prices and pathwise derivatives. code/firm/deephedge/firm_deephedge.py - Run
ml_hedge.hedging()(about a minute and a half on one core),surrogates(),calibration(),lm_check()andfig_hedge.py.
What to change next. Give the network the current variance as an input and add a variance swap as a second hedge; train on randomised Heston parameters and test on another model.
19.7 Build: deep hedging and learned pricers
Purpose. Hedges and prices learned where they save time or cost, checked against Book 5’s exact engines.
Interface. heston_paths; BandNet, hedge_pnl(policy, S, K, T, cost), entropic, expected_shortfall, train_hedge(S, K, T, cost, risk, seed, steps, batch, lr); bs_policy, ww_policy, no_hedge, band; SurrogateNet, fit_surrogate(X, y, dy, differential, seed), surrogate_values; fit_calibrator(vols, params, seed).
Rules. Every learned price, Greek and hedge is scored against an exact reference on data it was not trained on; training boxes and cost levels are recorded with the model.
Acceptance tests. code/firm/deephedge/tests/: the P&L of the delta hedge without costs matches the discrete hedging error’s scale; the entropic risk of a constant is minus the constant; the band network holds its position inside the band; Whalley–Wilmott’s band widens with cost; a differential surrogate of a known function recovers its slope; the calibrator inverts a simple map.
Stretch. A second hedging instrument; recurrent policies with the variance as a hidden state; arbitrage-free surrogate architectures.
Sources and further reading
- H. Buehler, L. Gonon, J. Teichmann and B. Wood, “Deep hedging”, Quantitative Finance 19(8), 2019.
- S. Imaki, K. Imajo, K. Ito, K. Minami and K. Nakagawa, “No-transaction band network”, arXiv:2103.01775, 2021.
- A. E. Whalley and P. Wilmott, “An asymptotic analysis of an optimal hedging model for option pricing with transaction costs”, Mathematical Finance 7(3), 1997.
- H. Föllmer and A. Schied, “Convex measures of risk and trading constraints”, Finance and Stochastics 6(4), 2002.
- B. Huge and A. Savine, “Differential machine learning”, arXiv:2005.02347, 2020.
- B. Horvath, A. Muguruza and M. Tomas, “Deep learning volatility”, Quantitative Finance 21(1), 2021.
- A. Hernandez, “Model calibration with neural networks”, SSRN 2812140, 2016.
19.8 Exercises
Exercise 19.1 ★
Show that the entropic risk measure is cash-invariant: . What is of a constant?
Solution
Solution of Exercise 19.1.
. For a constant, : holding cash needs more cash to be acceptable, that is, it frees .
Exercise 19.2 ★
Compute the Whalley–Wilmott half-width at the money half-way to maturity for , , and a Black–Scholes gamma of 0.0998.
Solution
Solution of Exercise 19.2.
: a band in delta, 0.229 wide, against the learned band’s 0.055 at the same cost.
Exercise 19.3 ★
Why is the indifference price above the Heston value (2.25) even without costs?
Solution
Solution of Exercise 19.3.
Daily hedging with one instrument cannot remove the risk of discrete rebalancing and of stochastic volatility, and the seller charges for the remaining risk with : 2.454 against 2.254, a risk premium of 0.20.
Exercise 19.4 ★★
Why does the deep hedge beat the Black–Scholes delta when there are no costs at all?
Solution
Solution of Exercise 19.4.
Under Heston with negative correlation, a fall in the spot comes with a rise in volatility, which raises the call’s value; the variance-minimising hedge ratio is the delta plus a correction for this co-movement (a smaller hedge for a call when ). The Black–Scholes delta at a fixed implied volatility ignores it; the network, trained on Heston paths, learns it.
Exercise 19.5 ★★
Why are the mean-reversion speed and the long-run variance hard for the calibration network, and what data would help?
Solution
Solution of Exercise 19.5.
Their effects on prices are nearly interchangeable over three maturities: faster mean reversion towards a long-run level and slower reversion towards a different level can produce similar term structures. Longer maturities, more of them, and variance-sensitive instruments (variance swaps) separate them.
Exercise 19.6 ★★
Find the flaw. “Our surrogate matches the pricer to 0.1% on the test set, so we use it for the stress scenarios.”
Solution
Solution of Exercise 19.6.
Stress scenarios lie outside the region the test set sampled, where the surrogate is extrapolating; its test error says nothing there. Validate on the scenarios themselves, or train on a box that contains them, and check Greeks as well as prices.
Exercise 19.7 ★★★
Coding. Derive the pathwise delta label for a digital call (). What goes wrong, and what does Huge and Savine’s approach do instead?
Solution
Solution of Exercise 19.7.
The pathwise derivative of with respect to is zero almost surely: the label says the delta is zero while the true delta is a density at the strike. Pathwise labels need a smooth payoff; the standard remedy is to smooth the discontinuity (replace the digital by a tight call spread) before differentiating, or to use likelihood-ratio labels.
Exercise 19.8 ★★★
Show that without costs and in a complete market (Black–Scholes, continuous trading) the entropic indifference price equals the Black–Scholes price for any .
Solution
Solution of Exercise 19.8.
In a complete market the claim is replicated exactly: a strategy has gains path by path. Any strategy for the seller can be written , so and, by cash invariance, . Taking the infimum over and subtracting the no-claim term leaves , whatever .
19.9 Problem: Hedging with Frictions
Problem 19.1
Weekend problem — when not to trade
The chapter’s call, hedges, surrogate and calibrator.
Part I — The problem.
- What is sold, how is it hedged, and what does the network see?
- What does daily delta hedging cost at 5 basis points, and how does that compare with the vega?
- Why a band network rather than a network that outputs the holding directly?
- What is the loss, and what does mean here?
Part II — Results.
- Give the indifference prices of the three hedges at each cost level.
- Compare the learned band with the Whalley–Wilmott band.
- What does training on expected shortfall change?
- Why does the Whalley–Wilmott hedge have the lowest mean cost at 10 basis points but not the lowest risk?
Part III — Surrogates and calibration.
- What do differential labels buy, and when?
- What errors does the calibration network make?
- How should the network and the optimiser be combined?
- Where would each learned pricer fail on a real desk?
Part IV — The verdict.
- State the named result: the indifference price as a function of the cost level, and the learned no-trade band’s width against the Whalley–Wilmott band.
- Would you quote the option at the indifference price? Why or why not?
- How would you validate a deep hedge before using it on a book?
- What would change with a second hedging instrument?
- How do you keep a surrogate honest when the model is recalibrated daily?
- What is the risk of training the hedge on one Heston parameter set?
- Where does automatic differentiation fit in this chapter?
- In one sentence: what does a hedging network learn that a delta does not?
Solution
Solution of Problem 19.1.
Part I.
- A one-month at-the-money call on 100 under Heston, hedged daily with the underlying; the network sees the time to maturity, the log-moneyness and, through the band, its current holding.
- 0.123 on average, more than one volatility point of vega (0.115).
- Not trading is then the default, the band is readable, and training is easier than learning to reproduce the previous holding exactly.
- The entropic risk of the P&L before the premium; is an absolute risk aversion per unit of price, so a loss of 1 is weighed times a loss of 0.
Part II.
- At 0, 5, 10, 20 basis points: deep 2.454, 2.558, 2.658, 2.840; Black–Scholes delta 2.475, 2.608, 2.743, 3.018; Whalley–Wilmott 2.475, 2.606, 2.712, 2.910.
- At the money half-way to maturity the learned band is 0.031 wide at 5 basis points and 0.104 at 20, against 0.181 and 0.288; it is narrower at the money and above it, and wider below the money.
- Lower expected shortfall (3.763 against 3.817) and higher entropic risk (2.671 against 2.658): each objective wins by its own measure.
- Its wide band trades least (mean cost 2.424), but the position strays further from the hedge, and the P&L’s standard deviation (0.703) and tails are the largest of the hedges.
Part III.
- Accuracy per sample: with 1 024 samples, price error 0.218 against 0.436 and delta error 0.021 against 0.059; at 4 096, equal prices and still better deltas.
- Parameter errors of 1.7% (initial variance) to 18.7% (mean-reversion speed) of each range, 13.2 basis points of spot in re-priced surfaces, 0.46 implied-volatility points on five surfaces.
- The network’s answer as the optimiser’s start (and as a check on its result).
- Outside the training box, on arbitrage constraints, and on market data that no model fits.
Part IV.
- Hedging with frictions. The indifference price rises from 2.454 without costs to 2.558, 2.658 and 2.840 at 5, 10 and 20 basis points; the learned band at the money is 0.031 and 0.104 wide at 5 and 20 basis points, against Whalley–Wilmott’s 0.181 and 0.288.
- As a bound on the price at which the desk is compensated for the risk it keeps, not as a quote: the market price and the book’s other positions matter.
- Out-of-sample paths, other models and parameters, stress paths, and a comparison with the asymptotic band and the delta on the book’s history.
- The network would also choose how much of the second instrument to hold, hedging volatility risk; the indifference price would fall towards the Heston value.
- Retrain or check it against the exact pricer after each recalibration, on a grid covering the new parameters.
- The hedge is optimal for one volatility dynamics; a different one (higher volatility of variance, a jump) can make it worse than the delta.
- It computes the network’s own derivatives in differential training, the hedge’s gradients in deep hedging, and exact Greeks of the reference pricer.
- When trading is worth its cost.
19.10 Interview questions
Interview question 19.1 ★ researcher, trader
Why does delta hedging with transaction costs call for a no-trade band?
Solution
Solution of Interview question 19.1.
Each rebalancing costs; small deviations from delta add little risk while trading them away costs a spread every time. The optimal policy leaves the position alone while it is close enough to delta and trades back to the band’s edge otherwise; the band widens with costs and narrows with gamma and risk aversion.
What the interviewer is looking for: the cost-risk trade-off and how the band depends on cost, gamma and risk aversion.
Interview question 19.2 ★★ researcher
What is deep hedging, and what are its inputs, outputs and loss?
Solution
Solution of Interview question 19.2.
A network policy trained on simulated paths: inputs are observable state (time, prices, holdings, other features), outputs are holdings in the hedging instruments, the loss is a convex risk measure of the terminal hedged P&L after costs.
What the interviewer is looking for: state, action, loss, and simulation as the training data.
Interview question 19.3 ★★ mle, researcher
Explain differential machine learning and why it helps with noisy Monte Carlo labels.
Solution
Solution of Interview question 19.3.
Train on labels and their derivatives with respect to the inputs, matching the network’s derivatives by automatic differentiation; pathwise derivative labels from Monte Carlo are nearly free and far less noisy relative to their signal, so fewer samples pin down the function and its slope.
What the interviewer is looking for: the loss and why derivative labels reduce the sample need.
Interview question 19.4 ★★ researcher
What is an indifference price, and how does it relate to a risk measure?
Solution
Solution of Interview question 19.4.
The premium that leaves the seller indifferent, under a utility or risk measure, between selling with the best hedge and not selling. With a cash-invariant risk measure it is the risk of the best hedged position without the premium.
What the interviewer is looking for: the definition and its link to cash invariance.
Interview question 19.5 ★★ mle
A calibration network is fast. Why not use it alone?
Solution
Solution of Interview question 19.5.
It is approximate, blind outside its training box, trained on model data that fit perfectly while market data do not, and gives no fit diagnostics; use it to start the optimiser and to check its result.
What the interviewer is looking for: approximation, extrapolation and model-data mismatch.
Interview question 19.6 ★★★ researcher
Your surrogate’s gamma turns negative for deep out-of-the-money strikes. What happened and how do you fix it?
Solution
Solution of Interview question 19.6.
The surrogate is extrapolating or under-trained where prices are tiny and its fit is judged by absolute error; nothing forces convexity. Train on log prices or implied volatilities, use differential labels including second derivatives, add convexity constraints or penalties, and sample the wings more densely.
What the interviewer is looking for: the cause and at least two remedies.