Quantitative Finance · Book 5 · Derivatives

Derivatives and Volatility

Derivatives and Volatility · Derivatives

12Rough Volatility and Forward-Variance Models

Take a long record of daily realised volatility for a stock index, and measure how far its logarithm moves over one day, two days and so on up to a month. Plot the mean squared move against the lag on log-log axes. The points lie on a line, and its slope, divided by two, estimates how smooth the path is. For a diffusion model of volatility, such as Heston’s, the answer at short lags is one half. In 2014 three researchers reported, for every index they measured, a number near 0.1: volatility is much rougher than the models desks used. The same number explains something the models of chapter 10 cannot do, which is the at-the-money skew of index options that keeps growing as expiry shrinks. This chapter starts with the object that makes such models practical, the forward-variance curve. It then covers Bergomi’s models of that curve, rough volatility and the rough Bergomi model, the path-dependent alternative, and the test every one of them faces, which is fitting index options and volatility-index options at once.

12.1 Forward variance

Chapter 10’s models describe the instantaneous variance vtv_t with a few parameters, and those parameters decide the term structure of at-the-money volatility. Desks prefer to take that term structure from the market and to model only how it moves. The object that makes this possible is the expected future variance, date by date.

Definition 12.1 (Forward variance)

The forward variance for date uu seen at time t≤ut\le u is ξt(u)=EtQ[vu]\xi_t(u)=\E^{\mathbb Q}_t[v_u], the pricing-measure expectation of the instantaneous variance at uu. The curve u↦ξt(u)u\mapsto\xi_t(u) is the forward-variance curve; ξ0\xi_0 is today’s.

The curve is read from prices. The expected integrated variance to TT is EQ[∫0Tvu du]=∫0Tξ0(u) du\E^{\mathbb Q}\bigl[\int_0^Tv_u\,du\bigr]=\int_0^T\xi_0(u)\,du. In a model where the price is a continuous process, that expectation is minus twice the price of the log contract, −2 EQ[ln⁡(ST/FT)]-2\,\E^{\mathbb Q}[\ln(S_T/F_T)], because Itô’s lemma gives ln⁡(ST/FT)=∫0Tvu dWu−12∫0Tvu du\ln(S_T/F_T)=\int_0^T\sqrt{v_u}\,dW_u-\frac12\int_0^Tv_u\,du for a driftless forward, and the stochastic integral has zero mean. The log payoff is a static strip of out-of-the-money options. With zero rates and k=ln⁡(K/F)k=\ln(K/F),

σVS2(T) T=∫0Tξ0(u) du=2∫−∞∞OTM(k) e−k dk,\sigma_{\mathrm{VS}}^2(T)\,T=\int_0^T\xi_0(u)\,du=2\int_{-\infty}^{\infty}\mathrm{OTM}(k)\,e^{-k}\,dk,

where OTM(k)\mathrm{OTM}(k) is the price of the put below the forward and of the call above it, per unit of forward. The left side is the fair strike of a variance swap (chapter 14 builds that product and its hedge), so σVS(T)\sigma_{\mathrm{VS}}(T) is called the variance-swap volatility. Its derivative in TT gives the curve: ξ0(T)=∂T(TσVS2(T))\xi_0(T)=\partial_T\bigl(T\sigma_{\mathrm{VS}}^2(T)\bigr).

Example 12.2 (The forward-variance curve of chapter 9’s surface)

Integrating each expiry’s strip on chapter 9’s surface gives variance-swap volatilities of 16.3%, 18.2%, 20.1%, 22.1% and 23.6% at one month, three months, six months, one year and two years, against at-the-money volatilities of 14.9%, 16.4%, 17.8%, 19.2% and 19.9%. With total variance linear between the pillars, the forward volatility ξ0\sqrt{\xi_0} is 16.3% up to one month, then 19.1%, 21.8%, 24.0% and, from one to two years, 24.9% (Figure 12.1).

The variance-swap volatility sits two to four points above the at-the-money volatility. The strip weights each option by e−ke^{-k}, that is by 1/K21/K^2 in strike, so the expensive low-strike puts of a negatively skewed smile count for more than the calls. The gap is a measure of the skew, and it grows with the expiry here because the wings widen. A forward-variance curve must be non-negative. A pillar where total variance falls is a calendar arbitrage: buy the longer variance swap and sell the shorter one. The build’s negative_intervals flags it before any model is calibrated.

The forward-variance curve of chapter 9’s surface. The variance-swap volatility comes from each expiry’s option strip and lies above the at-the-money volatility because the strip weights the skewed puts. Its forward is piecewise constant between the pillars. Data: the chapter’s code.
Figure 12.1. The forward-variance curve of chapter 9’s surface. The variance-swap volatility comes from each expiry’s option strip and lies above the at-the-money volatility because the strip weights the skewed puts. Its forward is piecewise constant between the pillars. Data: the chapter’s code.

12.2 Forward-variance models and Bergomi’s model

Definition 12.3 (Forward-variance model)

A forward-variance model specifies the dynamics of the whole curve (ξt(u))u≥t\bigl(\xi_t(u)\bigr)_{u\ge t}, starting from the market’s ξ0\xi_0, with each ξt(u)\xi_t(u) a martingale in tt under the pricing measure. The instantaneous variance is the curve’s short end, vt=ξt(t)v_t=\xi_t(t).

The martingale condition is not a modelling choice. ξt(u)\xi_t(u) is a conditional expectation of the same random variable vuv_u as tt moves, so it has no drift. Two things follow. The model reprices every variance swap by construction, whatever its parameters, because the initial curve is an input. And its parameters are left free to describe dynamics: how much each point of the curve moves, and how those moves correlate with the underlying and with each other.

Definition 12.4 (Bergomi model)

The Bergomi model is the forward-variance model with lognormal dynamics and exponentially decaying volatility along the curve. In its one-factor form,

dξt(u)=η e−κ(u−t) ξt(u) dWt1,ξt(u)=ξ0(u)exp⁡(ηe−κ(u−t)Xt−12η2e−2κ(u−t)E[Xt2]),d\xi_t(u)=\eta\,e^{-\kappa(u-t)}\,\xi_t(u)\,dW^1_t,\qquad \xi_t(u)=\xi_0(u)\exp\Bigl(\eta e^{-\kappa(u-t)}X_t-\tfrac12\eta^2e^{-2\kappa(u-t)}\E[X_t^2]\Bigr),

with Xt=∫0te−κ(t−s)dWs1X_t=\int_0^te^{-\kappa(t-s)}dW^1_s an Ornstein–Uhlenbeck factor and d⟨W1,WS⟩=ρ dtd\langle W^1,W^S\rangle=\rho\,dt for the Brownian motion WSW^S of the underlying. The two-factor form adds a second factor with a faster decay and mixes the two.

The exponential kernel ηe−κτ\eta e^{-\kappa\tau} is the instantaneous volatility of the forward variance τ=u−t\tau=u-t years ahead. Bergomi’s aim was to control separately three things that Heston ties together: the short forward skew, the spot–volatility correlation and the term structure of the volatility of volatility. Products whose value hangs on those quantities, such as reverse cliquets, Napoleons and options on realised variance, are the ones for which he designed it (Bergomi’s “Smile dynamics II”). The one-factor form is Markov in the two variables (St,Xt)(S_t,X_t), so it simulates as cheaply as Heston. The two-factor form fits the observed term structure of volatility of volatility better, at the cost of one more factor.

What an exponential kernel cannot do is explode. However fast κ\kappa is, the volatility of a forward variance a day ahead is at most η\eta. The short-dated skew of the model is driven by the volatility of short-dated forward variance, so it levels off like Heston’s (chapter 10). A second factor with a fast decay lifts the short end, but the power law of the market needs a kernel that is infinite at zero.

Left: the volatility of forward variance by horizon, for the rough kernel (H=0.1, =1.9) and a one-factor Bergomi kernel matched to it at one month and one year (=2.51, =1.08). Only the rough kernel explodes at short horizons. Right: one year of daily spot volatility from each model, driven by the same Brownian increments, with a flat 20% forward volatility. Data: the chapter’s code.
Figure 12.2. Left: the volatility of forward variance by horizon, for the rough kernel (H=0.1H=0.1, η=1.9\eta=1.9) and a one-factor Bergomi kernel matched to it at one month and one year (η=2.51\eta=2.51, κ=1.08\kappa=1.08). Only the rough kernel explodes at short horizons. Right: one year of daily spot volatility from each model, driven by the same Brownian increments, with a flat 20% forward volatility. Data: the chapter’s code.

12.3 Rough volatility

12.3.1 The evidence

Gatheral, Jaisson and Rosenbaum measured the smoothness of volatility directly. They took daily realised-variance estimates for many indices and computed the structure function m(q,Δ)=∣log⁡σt+Δ−log⁡σt∣q‾m(q,\Delta)=\overline{|\log\sigma_{t+\Delta}-\log\sigma_t|^q} for lags of one to thirty days. They found m(q,Δ)∝Δζqm(q,\Delta)\propto\Delta^{\zeta_q} with ζq≈Hq\zeta_q\approx Hq for every qq. That is the scaling of a fractional Brownian motion with Hurst exponent HH (One Quant Book 4, chapter 17), whose increments over a lag Δ\Delta have standard deviation proportional to ΔH\Delta^H. Their estimates were H=0.125H=0.125 for DAX futures, from one-hour windows of intraday data, and 0.1420.142 for the S&P 500, from daily realised variance. Across the 21 indices of the Oxford-Man realised library, ζ2/2\zeta_2/2 runs from 0.075 to 0.158 (Figure 12.3). A Brownian motion has H=12H=\frac12. A diffusion for volatility, Heston’s square-root process or an Ornstein–Uhlenbeck process for its logarithm, looks like a Brownian motion at short lags, so it has H=12H=\frac12 too.

Definition 12.5 (Rough volatility)

Rough volatility is the property, and the class of models built on it, that the logarithm of instantaneous volatility behaves at short time scales like a fractional Brownian motion with Hurst exponent H<12H<\frac12, typically near 0.1. Its paths are rougher than any diffusion’s: the typical move over a lag Δ\Delta scales like ΔH\Delta^H rather than Δ1/2\Delta^{1/2}.

Left: the structure function of simulated daily log-volatility over eight years. The slope on log-log axes is 2H. The rough path has H=0.1; the diffusion is an Ornstein–Uhlenbeck process; the third series is the diffusion seen through a realised-variance estimate with a 0.08 error in log. Right: the published estimates for 21 indices. Data: the chapter’s code; Gatheral, Jaisson and Rosenbaum (2014), table B.1.
Figure 12.3. Left: the structure function of simulated daily log-volatility over eight years. The slope on log-log axes is 2H2H. The rough path has H=0.1H=0.1; the diffusion is an Ornstein–Uhlenbeck process; the third series is the diffusion seen through a realised-variance estimate with a 0.08 error in log. Right: the published estimates for 21 indices. Data: the chapter’s code; Gatheral, Jaisson and Rosenbaum (2014), table B.1.

The regression is simple, and its reading needs care. Realised variance is itself an estimate. From 78 five-minute returns its relative standard error is about 2/78=0.16\sqrt{2/78}=0.16, so the logarithm of realised volatility carries an error near 0.08 each day. The error is independent from day to day, so it adds a constant to m(2,Δ)m(2,\Delta) at every lag and flattens the log-log line. On simulated data (Figure 12.3), a diffusion whose true HH is 12\frac12, read through that error, returns 0.17. Averaging over a window works the other way: Gatheral and his coauthors simulated their own rough model with H=0.14H=0.14 and recovered 0.16 to 0.18 from realised-variance proxies. The case for roughness therefore rests on more than one regression. The estimate holds across indices, across moments qq, and across the estimators of the paper, and it has a consequence in option prices, where no measurement error enters.

12.3.2 The skew that does not flatten

Define the at-the-money skew ψ(T)=∂σBS(k,T)/∂k\psi(T)=\partial\sigma_{\mathrm{BS}}(k,T)/\partial k at k=0k=0. On the S&P 500 on 20 June 2013, Gatheral, Jaisson and Rosenbaum fitted ∣ψ(T)∣≈A T−0.4|\psi(T)|\approx A\,T^{-0.4} across expiries. Diffusive stochastic volatility models such as Heston, SABR and one-factor Bergomi give a skew that tends to a constant as T→0T\to0 and decays like a sum of exponentials for longer expiries. Fukasawa showed that a volatility driven by a fractional Brownian motion gives ψ(T)∼TH−1/2\psi(T)\sim T^{H-1/2} for short expiries. With H=0.1H=0.1 the exponent is −0.4-0.4, the market’s. One parameter explains both the roughness of the time series and the shape of the skew.

Definition 12.6 (Rough Bergomi model)

The rough Bergomi model is the forward-variance model with the power-law kernel. Its spot variance is

vt=ξ0(t)exp⁡(ηYt−12η2t2H),Yt=2H∫0t(t−s)H−1/2 dWs1,v_t=\xi_0(t)\exp\Bigl(\eta Y_t-\tfrac12\eta^2t^{2H}\Bigr),\qquad Y_t=\sqrt{2H}\int_0^t(t-s)^{H-1/2}\,dW^1_s,

and the underlying follows dSt/St=vt (ρ dWt1+1−ρ2 dWt⊥)dS_t/S_t=\sqrt{v_t}\,\bigl(\rho\,dW^1_t+\sqrt{1-\rho^2}\,dW^\perp_t\bigr). YY is a Riemann–Liouville fractional Brownian motion with Var⁡(Yt)=t2H\Var(Y_t)=t^{2H}, so E[vt]=ξ0(t)\E[v_t]=\xi_0(t): the model fits the forward-variance curve by construction. It has three parameters, HH, η\eta and ρ\rho.

The forward variances of the model are ξt(u)=ξ0(u)exp⁡(η2H∫0t(u−s)H−1/2dWs1−12η2(u2H−(u−t)2H))\xi_t(u)=\xi_0(u)\exp\bigl(\eta\sqrt{2H}\int_0^t(u-s)^{H-1/2}dW^1_s-\frac12\eta^2(u^{2H}-(u-t)^{2H})\bigr), so the volatility of ξt(u)\xi_t(u) is η2H (u−t)H−1/2\eta\sqrt{2H}\,(u-t)^{H-1/2}. That is the power-law kernel of Figure 12.2. With H=0.1H=0.1 and η=1.9\eta=1.9 it is 0.85 at one year, 2.30 at one month, 4.13 at one week and 7.76 at one day, three times the matched Bergomi kernel. The spot volatility path is correspondingly wild: its logarithm moves by 0.53 on an average day in the right panel of the figure, against 0.06 for the Bergomi path. The price of this realism is that the model is not Markov. YtY_t depends on the whole past of W1W^1 with weights that decay slowly, so no finite set of state variables summarises it. Monte Carlo (One Quant Book 4, chapter 26) is the only general pricing method, and a PDE is not available.

Proposition 12.7 (Exact simulation on a grid)

On a grid t1<⋯<tnt_1<\dots<t_n the vector (Yt1,…,Ytn,Wt11,…,Wtn1)(Y_{t_1},\dots,Y_{t_n},W^1_{t_1},\dots,W^1_{t_n}) is Gaussian with, for u≤vu\le v,

Cov⁡(Yu,Yv)=u2HG(v/u),G(x)=2H∫01(1−s)H−12(x−s)H−12 ds,\Cov(Y_u,Y_v)=u^{2H}G(v/u),\qquad G(x)=2H\int_0^1(1-s)^{H-\frac12}(x-s)^{H-\frac12}\,ds,

and, for any uu and vv,

Cov⁡(Yv,Wu1)=2HH+12(vH+12−(v−min⁡(u,v))H+12),\Cov(Y_v,W^1_u)=\frac{\sqrt{2H}}{H+\frac12}\Bigl(v^{H+\frac12}-\bigl(v-\min(u,v)\bigr)^{H+\frac12}\Bigr),

so one Cholesky factorisation of this 2n×2n2n\times2n matrix gives exact samples of both at the grid dates.

Proof. Both are Wiener integrals against the same W1W^1, so the vector is Gaussian and its covariances are integrals of products of kernels. For u≤vu\le v, Cov⁡(Yu,Yv)=2H∫0u(u−s)H−12(v−s)H−12ds\Cov(Y_u,Y_v)=2H\int_0^u(u-s)^{H-\frac12}(v-s)^{H-\frac12}ds; substituting s=us′s=us' gives u2HG(v/u)u^{2H}G(v/u). The cross term is 2H∫0min⁡(u,v)(v−s)H−12ds\sqrt{2H}\int_0^{\min(u,v)}(v-s)^{H-\frac12}ds, integrated directly. ∎

The factorisation costs O(n3)O(n^3) once and each path O(n2)O(n^2). That is cheap for the hundred steps the tutorial uses and expensive for thousands of steps. The hybrid scheme of Bennedsen, Lunde and Pakkanen treats the kernel exactly near zero, where it is singular, and as a step function further away. The far part is a convolution, computed with a fast Fourier transform, which brings the cost to O(nlog⁡n)O(n\log n) per path. Prices then come from a second saving, the mixing formula of chapter 10. Given the path of W1W^1, the logarithm of STS_T is normal with mean ρ∫v dW1−12∫v dt\rho\int\sqrt v\,dW^1-\frac12\int v\,dt and variance (1−ρ2)∫v dt(1-\rho^2)\int v\,dt. Each path therefore contributes a Black price rather than a payoff, and the noise of the skew estimate falls sharply.

Example 12.8 (The skew term structure)

Rough Bergomi with a flat 20% forward volatility, H=0.1H=0.1, η=1.9\eta=1.9 and ρ=−0.9\rho=-0.9 (illustrative values, not a calibration) gives an at-the-money skew of −1.76-1.76 at one week and −0.32-0.32 at one year. A power law fitted from one week to two years has exponent −0.44-0.44. On the same expiries chapter 9’s surface has −1.44-1.44 and −0.25-0.25 and exponent −0.44-0.44 as well. Heston calibrated to that surface in chapter 10 comes within 0.04 of it from three months on, but its skew at one week is −0.70-0.70, and its exponent over the first two months is −0.09-0.09: nearly flat (Figure 12.4).

The at-the-money skew by expiry on log-log axes, from one week to two years. The surface and rough Bergomi follow straight lines; Heston’s skew levels off below two months. Data: the tutorial.
Figure 12.4. The at-the-money skew by expiry on log-log axes, from one week to two years. The surface and rough Bergomi follow straight lines; Heston’s skew levels off below two months. Data: the tutorial.

The rough Bergomi fit also carries a warning. Its HH is 0.1, so the short-expiry rule ψ∼TH−1/2\psi\sim T^{H-1/2} predicts an exponent of −0.40-0.40. The fit over one week to two years returns −0.44-0.44, because the rule is a limit and two years is not short. Reading H=12+αH=\frac12+\alpha off a fitted exponent α\alpha is therefore biased by a few hundredths. The surface’s −0.44-0.44 reads as H=0.06H=0.06 by the rule, but as about 0.1 once the same fit on the model is used as the yardstick. This is the chapter’s weekend problem.

12.4 Path-dependent volatility

Roughness is one reading of the data. Guyon and Lekeufack proposed another in 2023. Volatility, in their account, is mostly a function of the path the price has already taken.

Definition 12.9 (Path-dependent volatility model)

A path-dependent volatility model makes the instantaneous volatility a function of past returns. In the Guyon–Lekeufack form it is

σt=β0+β1R1,t+β2R2,t,R1,t=∫−∞tK1(t−s) dSsSs,R2,t=∫−∞tK2(t−s) (dSsSs)2 ⁣/dt,\sigma_t=\beta_0+\beta_1R_{1,t}+\beta_2\sqrt{R_{2,t}},\qquad R_{1,t}=\int_{-\infty}^tK_1(t-s)\,\frac{dS_s}{S_s},\qquad R_{2,t}=\int_{-\infty}^tK_2(t-s)\,\Bigl(\frac{dS_s}{S_s}\Bigr)^2\!\Big/dt,

a trend feature R1R_1 (a weighted sum of past returns) and an activity feature R2R_2 (a weighted sum of past squared returns), with decaying kernels K1,K2K_1,K_2 that integrate to one and β1<0\beta_1<0.

Guyon and Lekeufack report that past returns explain up to 90% of the variance of index implied volatility and 65% of that of realised volatility. With exponential kernels the features are Markov: two exponentials per kernel give a four-factor model, and Gazzani and Guyon show that it fits S&P 500 and VIX smiles jointly. The model captures what traders call the leverage effect, and more: the volatility responds to the size of recent moves and, asymmetrically, to their sign.

Example 12.10 (One large day)

Take β0=0.04\beta_0=0.04, β1=−0.06\beta_1=-0.06, β2=0.65\beta_2=0.65 and exponential kernels with rates κ1=25\kappa_1=25 and κ2=15\kappa_2=15 a year (illustrative values). With no trend the calm level is the fixed point σ∗=β0/(1−β2)=11.4%\sigma^*=\beta_0/(1-\beta_2)=11.4\%. A day of −4%-4\% lifts volatility to 22.0%. The excess halves in ten trading days, and a month (21 trading days) later volatility is still 13.8%. A day of +4%+4\% first lowers volatility, to 10.6%, because the trend term dominates. The activity term decays more slowly, so volatility then drifts up to 12.4% before settling (Figure 12.5, left).

Rough and path-dependent models are rivals only in part. A path-dependent volatility has no randomness of its own: it is driven by the price’s shocks, so over an instant volatility and price move with correlation −1-1, and its trend kernel does the work that ρ\rho does in the other models. The practical differences are that the path-dependent model is Markov in a few factors and that its parameters come from time series as well as from options.

12.5 Joint calibration to index and volatility-index options

The volatility index (One Quant Book 1, chapter 25) is computed from a strip of 30-day index options, the same strip that gives the variance-swap volatility. Idealised, its square is the average forward variance over the next thirty days:

VIXT2=1Δ∫TT+ΔξT(u) du,Δ=30365.\mathrm{VIX}_T^2=\frac1\Delta\int_T^{T+\Delta}\xi_T(u)\,du,\qquad\Delta=\tfrac{30}{365}.

A forward-variance model therefore prices VIX futures and options. They are the most direct test of its dynamics, because they are claims on the curve itself. Two facts follow at once. The VIX future for date TT is EQ[VIXT]\E^{\mathbb Q}[\mathrm{VIX}_T], while the forward variance-swap volatility is EQ[VIXT2]\sqrt{\E^{\mathbb Q}[\mathrm{VIX}_T^2]}, so the future is below the forward volatility by a convexity term that grows with the volatility of the VIX. And a model with a given kernel produces one VIX smile, with no free parameter left to shape it once the index smile has fixed HH, η\eta and ρ\rho.

Example 12.11 (The VIX in rough Bergomi)

With the parameters of Example 12.8, the one-month VIX has E[VIX2]=0.0400\E[\mathrm{VIX}^2]=0.0400, the forward variance, as it must. The future is 18.74 against a forward variance-swap volatility of 20: a convexity gap of 1.26 points. The implied volatility of one-month VIX options is 123.1% at a strike of 80% of the future, 123.9% at the money and 125.7% at 160%. The model’s VIX is nearly lognormal, and its smile is nearly flat, with a level set by η\eta and HH. Guerreiro and Guerra make the same observation, and contrast it with the upward-sloping VIX smiles of the market.

Left: the response of the path-dependent volatility of  to one large day, down or up; volatility reacts asymmetrically to the sign. Right: the smile of one-month VIX options in rough Bergomi with the parameters that give the skew of : nearly flat. Data: the chapter’s code.
Figure 12.5. Left: the response of the path-dependent volatility of Example 12.10 to one large day, down or up; volatility reacts asymmetrically to the sign. Right: the smile of one-month VIX options in rough Bergomi with the parameters that give the skew of Figure 12.4: nearly flat. Data: the chapter’s code.

Fitting the S&P 500 smile and the VIX smile with one model is the joint calibration problem. It is a sharp test because the two markets constrain the same object from two sides. The index smile fixes the skew and the volatility of short-dated variance, and the VIX smile fixes the distribution of that variance directly. Guyon conjectured that no model with continuous paths could fit both. Gatheral, Jusselin and Rosenbaum answered in 2020 with the quadratic rough Heston model, which adds to roughness a feedback of past returns on volatility, the same ingredient as the path-dependent models. Guyon later solved it exactly with a non-parametric construction, a martingale Schrödinger problem that builds a joint law fitting both smiles, and the four-factor path-dependent model fits both with a few parameters. For a desk the lesson is practical. Products that pay on the VIX, and index products whose value depends on the volatility of variance, such as cliquets and options on realised variance, need a model tested against VIX options as well as index options. A model fitted to the index smile alone is untested on exactly the dynamics those products pay for.

12.6 Tutorial: rough volatility from the time series to the skew

Goal. Estimate roughness from a volatility path, read the forward-variance curve from a surface, simulate rough Bergomi exactly, and fit the skew term structure. End state: the five figures of the chapter and the numbers of the weekend problem.

  1. The roughness estimator. The structure function and the log-log slope; run it on fbm(1500, 0.3, seed) and check that it returns about 0.3:

    def structure_function(x: np.ndarray, lags, q: float = 2.0) -> np.ndarray:
        """m(q, lag) = mean |x_{t + lag} - x_t|^q."""
        return np.array([np.mean(np.abs(x[lag:] - x[:-lag]) ** q) for lag in lags])
    
    
    def roughness(x: np.ndarray, lags=range(1, 31), q: float = 2.0) -> float:
        """Estimate of H: the slope of log m(q, lag) on log lag, divided by q."""
        lags = np.asarray(list(lags), float)
        slope = np.polyfit(np.log(lags), np.log(structure_function(x, lags.astype(int), q)), 1)[0]
        return float(slope / q)
    Listing 12.1. Structure function and roughness estimate. code/firm/fwdvar/firm_fwdvar.py
  2. The Gaussian skeleton. The joint covariance of the Volterra process and its Brownian motion, from Proposition 12.7:

    def joint_cov(times, hurst: float) -> np.ndarray:
        """Covariance of (Y_{t_1..t_n}, W1_{t_1..t_n}), a 2n x 2n matrix."""
        t = np.asarray(times, float)
        lo, hi = np.minimum.outer(t, t), np.maximum.outer(t, t)
        cyy = lo ** (2 * hurst) * _g_ratio(hi / lo, hurst)
        c = math.sqrt(2 * hurst) / (hurst + 0.5)
        cyw = c * (t[:, None] ** (hurst + 0.5) - (t[:, None] - lo) ** (hurst + 0.5))   # Cov(Y_{t_i}, W_{t_j})
        return np.block([[cyy, cyw], [cyw.T, lo]])
    Listing 12.2. Covariance of (Y,W1)(Y,W^1) on a grid. code/firm/fwdvar/firm_fwdvar.py
  3. Paths and prices. Simulate on a grid with antithetic pairs, accumulate the Itô sum and the integrated variance, and price by the mixing formula:

    def rbergomi_integrals(xi0, hurst: float, eta: float, t: float, steps: int, n_paths: int,
                           seed: int) -> tuple[np.ndarray, np.ndarray]:
        """Per path, the Ito sum I = sum sqrt(v_i) dW1_i and the integrated variance Q = sum v_i dt on a uniform
        grid of `steps` steps to t (left points; antithetic pairs). xi0 is a number or a function of time."""
        xi = xi0 if callable(xi0) else (lambda u: float(xi0))
        grid = np.linspace(t / steps, t, steps)
        low = _chol(joint_cov(grid, hurst))
        rng = np.random.default_rng(seed)
        half = rng.standard_normal((n_paths // 2, 2 * steps))
        g = np.vstack([half, -half]) @ low.T
        y, w = g[:, :steps], g[:, steps:]
        left = np.concatenate([[0.0], grid[:-1]])
        xis = np.array([xi(u) for u in left])
        v = np.empty((len(g), steps))
        v[:, 0] = xis[0]
        v[:, 1:] = xis[1:] * np.exp(eta * y[:, :-1] - 0.5 * eta ** 2 * left[1:] ** (2 * hurst))
        dw = np.diff(np.concatenate([np.zeros((len(g), 1)), w], axis=1), axis=1)
        return np.sum(np.sqrt(v) * dw, axis=1), np.sum(v, axis=1) * (t / steps)
    
    
    def mixing_calls(i_sum: np.ndarray, q: np.ndarray, rho: float, strikes) -> np.ndarray:
        """Call prices on a forward of 1 by the mixing formula: given the volatility driver, ln S_T is normal
        with mean rho I - Q / 2 and variance (1 - rho^2) Q, so the call is a Black price averaged over paths."""
        s1 = np.exp(rho * i_sum - 0.5 * rho * rho * q)
        sd = np.sqrt((1 - rho * rho) * q)
        out = []
        for k in strikes:
            d1 = (np.log(s1 / k) + 0.5 * sd * sd) / sd
            out.append(float(np.mean(s1 * ncdf(d1) - k * ncdf(d1 - sd))))
        return np.array(out)
    Listing 12.3. Rough Bergomi by exact simulation and the mixing formula. code/firm/fwdvar/firm_fwdvar.py
  4. Run dv_roughvol.roughness_table(), variance_curve(), skew_term(), vix_smile() and fig_roughvol.py.

What to change next. Replace the fixed 100 steps by 25 and 400 and watch the one-week skew; set H=0.3H=0.3 and refit the exponent; add a noise of 0.04 instead of 0.08 to the diffusion’s log-volatility and re-estimate HH.

12.7 Build: forward variance and rough Bergomi

Purpose. The miniature firm’s forward-variance curve, which every volatility product of Part III reads (variance swaps, VIX futures, cliquets), and a rough-volatility engine for short-dated index options.

Interface. log_strip_variance(vol_of_k, t); ForwardVarianceCurve.from_vs_vols(times, vols) with xi(u), total_variance(t), vs_vol(t), forward_vol(t1, t2), negative_intervals(); joint_cov(times, H); rbergomi_integrals(xi0, H, eta, t, steps, n_paths, seed); mixing_calls; rbergomi_smile; rbergomi_vix; atm_skew; power_law_fit; fbm; structure_function; roughness.

Rules. Total variance is non-decreasing, or the curve is rejected before any model sees it; E[vt]=ξ0(t)\E[v_t]=\xi_0(t) exactly in the simulator (the variance’s compensator uses the same t2Ht^{2H} as the covariance); simulated prices are martingale-checked (the forward is repriced within its standard error); seeds are explicit.

Acceptance tests. code/firm/fwdvar/tests/: a flat smile’s strip returns its variance; the curve’s pieces add up; G(1)=1G(1)=1 and the covariance reduces to the Brownian one at H=12H=\frac12; with η=0\eta=0 and ρ=0\rho=0 the smile is flat at ξ0\sqrt{\xi_0}; a negative correlation gives a negative skew; E[VIX2]=ξ0\E[\mathrm{VIX}^2]=\xi_0; the estimator recovers HH from exact fractional Brownian paths.

Stretch. The hybrid scheme with a fast Fourier transform, for grids of thousands of steps; a two-factor Bergomi model with the same interface; the four-factor path-dependent model and its VIX.

Sources and further reading

  • J. Gatheral, T. Jaisson and M. Rosenbaum, “Volatility is rough”, Quantitative Finance 18(6) (2018) 933–949; arXiv 1410.3394.
  • C. Bayer, P. Friz and J. Gatheral, “Pricing under rough volatility”, Quantitative Finance 16(6) (2016) 887–904.
  • M. Bennedsen, A. Lunde and M. S. Pakkanen, “Hybrid scheme for Brownian semistationary processes”, Finance and Stochastics 21(4) (2017) 931–965.
  • L. Bergomi, “Smile dynamics II”, SSRN 1493302.
  • J. Guyon and J. Lekeufack, “Volatility is (mostly) path-dependent”, Quantitative Finance 23(9) (2023) 1221–1258.
  • G. Gazzani and J. Guyon, “Pricing and calibration in the 4-factor path-dependent volatility model”, arXiv 2406.02319 (2024).
  • J. Gatheral, P. Jusselin and M. Rosenbaum, “The quadratic rough Heston model and the joint S&P 500/VIX smile calibration problem”, arXiv 2001.01789 (2020).
  • J. Guyon, “Dispersion-constrained martingale Schrödinger problems and the exact joint S&P 500/VIX smile calibration puzzle”, Finance and Stochastics (2023).

12.8 Exercises

Exercise 12.1 ★

The one-year and two-year variance-swap volatilities are 20% and 22%. Give the forward volatility between one and two years.

Solution

Solution of Exercise 12.1.

Total variances are 0.040.04 and 2×0.0484=0.09682\times0.0484=0.0968; the forward variance is 0.05680.0568 a year, a forward volatility of 23.8%.

Exercise 12.2 ★

Why is the variance-swap volatility above the at-the-money volatility when the smile is skewed to the downside?

Solution

Solution of Exercise 12.2.

The strip weights each out-of-the-money option by 1/K21/K^2, so low strikes count more than high ones. With a downside skew the low-strike puts are priced at higher volatilities than the at-the-money option, and the weighted average of variance is above the at-the-money variance.

Exercise 12.3 ★

Doubling the lag multiplies m(2,Δ)m(2,\Delta) by 1.15. Estimate HH.

Solution

Solution of Exercise 12.3.

m(2,Δ)∝Δ2Hm(2,\Delta)\propto\Delta^{2H}, so 22H=1.152^{2H}=1.15 and H=ln⁡1.15/(2ln⁡2)=0.10H=\ln1.15/(2\ln2)=0.10.

Exercise 12.4 ★★

Show that in rough Bergomi E[vt]=ξ0(t)\E[v_t]=\xi_0(t) and that the volatility of ξt(u)\xi_t(u) is η2H (u−t)H−1/2\eta\sqrt{2H}\,(u-t)^{H-1/2}.

Solution

Solution of Exercise 12.4.

YtY_t is Gaussian with mean zero and variance 2H∫0t(t−s)2H−1ds=t2H2H\int_0^t(t-s)^{2H-1}ds=t^{2H}, so E[eηYt]=eη2t2H/2\E[e^{\eta Y_t}]=e^{\eta^2t^{2H}/2} and E[vt]=ξ0(t)\E[v_t]=\xi_0(t). For u≥tu\ge t, ξt(u)=Et[vu]\xi_t(u)=\E_t[v_u]: split YuY_u into 2H∫0t(u−s)H−1/2dWs\sqrt{2H}\int_0^t(u-s)^{H-1/2}dW_s, known at tt, and an independent remainder of variance (u−t)2H(u-t)^{2H}; taking the expectation of the remainder’s exponential gives the formula of the text. Its only tt-dependence through WW is η2H∫0t(u−s)H−1/2dWs\eta\sqrt{2H}\int_0^t(u-s)^{H-1/2}dW_s, whose differential is η2H(u−t)H−1/2dWt\eta\sqrt{2H}(u-t)^{H-1/2}dW_t: that is the volatility of ξt(u)\xi_t(u).

Exercise 12.5 ★★

In Example 12.11 the VIX future is 18.74 while the forward variance-swap volatility is 20. Explain the gap, and say what a larger η\eta would do to it.

Solution

Solution of Exercise 12.5.

The future is E[VIX]\E[\mathrm{VIX}] and the forward volatility E[VIX2]\sqrt{\E[\mathrm{VIX}^2]}; by Jensen’s inequality E[VIX]≤E[VIX2]\E[\mathrm{VIX}]\le\sqrt{\E[\mathrm{VIX}^2]}, with a gap that grows with the variance of the VIX. A larger η\eta makes the VIX more volatile and widens the gap: with η=2.5\eta=2.5 the future is 17.88.

Exercise 12.6 ★★

In Example 12.10, derive the calm level σ∗\sigma^* and explain why volatility first falls after a +4%+4\% day and then rises above σ∗\sigma^*.

Solution

Solution of Exercise 12.6.

In the calm state R1=0R_1=0 and R2=σ2R_2=\sigma^2, so σ∗=β0+β2σ∗\sigma^*=\beta_0+\beta_2\sigma^* and σ∗=0.04/0.35=11.4%\sigma^*=0.04/0.35=11.4\%. After a +4%+4\% day, β1R1\beta_1R_1 is negative and larger in size than the rise of β2R2\beta_2\sqrt{R_2}, so volatility falls; R1R_1 decays at rate 25 and R2R_2 at 15, so the activity term outlasts the trend term and volatility rises above σ∗\sigma^*, to 12.4%, before both fade.

Exercise 12.7 ★★★

Coding. Estimate HH on the chapter’s diffusion path (Ornstein–Uhlenbeck log-volatility) without noise and with realised-variance noise of 0.08 in log, over five seeds. What range do you find?

Solution

Solution of Exercise 12.7.

Over five seeds: 0.45 to 0.50 without noise, the diffusion’s 12\frac12 lowered a little by mean reversion; 0.20 to 0.23 with noise. The rough path of the same code reads 0.10 to 0.12.

Exercise 12.8 ★★★

Find the flaw. “The structure function of our daily realised volatility gives H=0.17H=0.17, far below one half, so the desk’s diffusive volatility model is refuted.”

Solution

Solution of Exercise 12.8.

Daily realised volatility is measured with an error of about 0.08 in log, independent from day to day. The error adds a constant to the structure function at every lag and flattens its slope: a diffusion with H=12H=\frac12 read through it gives about 0.17 to 0.23. The estimate refutes nothing until the noise is modelled, for instance by subtracting its contribution, twice its variance, from m(2,Δ)m(2,\Delta) before the regression; the evidence from the skew term structure, where no measurement noise enters, is the stronger test.

12.9 Problem: The Skew That Will Not Flatten

Problem 12.1

Weekend problem — the power law of the at-the-money skew

An index desk prices short-dated options with the Heston model of chapter 10 and wants to know whether a rough model would serve it better. It has chapter 9’s surface, the Heston calibration, and the chapter’s rough Bergomi simulator.

Part I — The curve.

  1. Give the one-year variance-swap and at-the-money volatilities of the surface.
  2. Give the forward volatility between six months and one year.
  3. If the one-year variance-swap volatility rises one point and the six-month one does not move, what does that forward volatility become?
  4. Give the variance-swap strike, in volatility, of a variance swap starting in one year and ending in two.
  5. Is the curve free of calendar arbitrage?

Part II — Kernels.

  1. Give the rough kernel’s volatility of forward variance at one week, one month and one year (H=0.1H=0.1, η=1.9\eta=1.9).
  2. Give the one-factor Bergomi kernel that matches it at one month and one year.
  3. By what factor does the rough kernel exceed it one day ahead?
  4. Why can no choice of κ\kappa make the exponential kernel match the rough one at every horizon?
  5. What would a second, fast Bergomi factor do?

Part III — Skews.

  1. Give the surface’s at-the-money skew at one week and one year.
  2. Give Heston’s at the same expiries.
  3. Give the power-law exponents fitted from one week to two years for the surface and for rough Bergomi.
  4. Give Heston’s exponent over the first two months.
  5. Check the surface’s exponent against the ratio of its one-week and one-year skews.

Part IV — Judgement.

  1. What HH does the short-expiry rule ψ∼TH−1/2\psi\sim T^{H-1/2} read from the surface’s exponent?
  2. Rough Bergomi with H=0.1H=0.1 gives the same fitted exponent. What does that say about the reading?
  3. Compare with the HH estimated from index time series.
  4. State the named result: the power-law exponent of the at-the-money skew term structure fitted on the chapter’s surface, and the Hurst exponent it implies.
  5. Would you move the desk’s short-dated pricing to rough Bergomi? Name one cost.
Solution

Solution of Problem 12.1.

1. 22.1% and 19.2%. 2. 24.0%. 3. 25.8%. 4. 24.9%. 5. Yes: total variance rises at every pillar, and negative_intervals() is empty. 6. 4.13, 2.30 and 0.85. 7. η=2.51\eta=2.51, κ=1.08\kappa=1.08. 8. 3.1 (7.76 against 2.50). 9. The exponential is at most η\eta at zero horizon; the power law is infinite there. Two points fix η\eta and κ\kappa, and the curves then differ everywhere else. 10. It adds a fast-decaying term that lifts the short end; a sum of exponentials can imitate the power law over a chosen range of horizons, never down to zero. 11. −1.44-1.44 and −0.25-0.25. 12. −0.70-0.70 and −0.25-0.25. 13. −0.44-0.44 for both. 14. −0.09-0.09: nearly flat. 15. 1.44/0.25=5.71.44/0.25=5.7, and a power law with exponent −0.44-0.44 predicts 520.44=5.752^{0.44}=5.7 between one week and one year. 16. H=12−0.44=0.06H=\frac12-0.44=0.06. 17. The rule applied to a model with a known H=0.1H=0.1 returns 0.06 as well: over this range of expiries it reads low by about 0.04, so the surface is consistent with HH near 0.1. 18. Time series give 0.075 to 0.158 across 21 indices, 0.124 for the S&P 500 in the same table: consistent with 0.1. 19. Exponent −0.44-0.44; H=0.06H=0.06 by the short-expiry rule, about 0.1 once the rule is calibrated on rough Bergomi. 20. Yes for short-dated index options, whose skew Heston cannot produce. Costs: Monte Carlo only (no PDE, slower and noisier Greeks), and the VIX smile the model implies must be checked against the market’s before its volatility-of-volatility products are trusted.

12.10 Interview questions

Interview question 12.1 ★ trader, researcher

What is forward variance, and how do you read it from the market?

Solution

Solution of Interview question 12.1.

ξt(u)=EtQ[vu]\xi_t(u)=\E^{\mathbb Q}_t[v_u], the expected instantaneous variance at a future date. The expected integrated variance to TT is minus twice the log contract’s price, which a strip of out-of-the-money options replicates; that is the variance-swap strike TσVS2(T)T\sigma^2_{\mathrm{VS}}(T), and its derivative in TT is ξ0(T)\xi_0(T).

What the interviewer is looking for: the definition, the log-contract strip, and the derivative in expiry.

Interview question 12.2 ★ researcher

What does “volatility is rough” mean, and what is the evidence?

Solution

Solution of Interview question 12.2.

The logarithm of volatility behaves at short time scales like a fractional Brownian motion with HH near 0.1, far rougher than a diffusion’s 12\frac12. Evidence: the scaling of the structure function of realised volatility across many indices and moments, and the power-law at-the-money skew ψ∼TH−1/2\psi\sim T^{H-1/2} in option prices. Caveat: the time-series estimate is sensitive to measurement noise.

What the interviewer is looking for: both kinds of evidence, and the caveat.

Interview question 12.3 ★★ researcher, trader

Why does Heston fail to fit the short-dated skew, and how does rough volatility fix it?

Solution

Solution of Interview question 12.3.

In a diffusive stochastic volatility model the short skew comes from the volatility of variance over a short time, which is bounded, so the skew tends to a constant as T→0T\to0. A rough volatility has a volatility of forward variance that explodes like τH−1/2\tau^{H-1/2} at short horizons, producing a skew ∼TH−1/2\sim T^{H-1/2} that keeps growing, like the market’s.

What the interviewer is looking for: the short-expiry limit, and the kernel.

Interview question 12.4 ★★ developer

How would you simulate the rough Bergomi model, and what does it cost?

Solution

Solution of Interview question 12.4.

The model is not Markov, so Monte Carlo. Exact: the Volterra process and its Brownian motion are jointly Gaussian on the grid, so Cholesky of the 2n×2n2n\times2n covariance, O(n3)O(n^3) once and O(n2)O(n^2) per path. Faster: the hybrid scheme, exact near the singularity and a fast-Fourier convolution elsewhere, O(nlog⁡n)O(n\log n). Price with the mixing formula (conditioning on the variance driver) to cut noise; check the martingale property.

What the interviewer is looking for: non-Markov, the Gaussian structure, the costs, variance reduction.

Interview question 12.5 ★★ researcher, risk

What is the joint S&P 500/VIX calibration problem, and why does a desk care?

Solution

Solution of Interview question 12.5.

Fitting the index smile and the VIX smile with one model. The index smile constrains the skew and the volatility of short-dated variance; the VIX smile constrains the law of that variance directly, and simple models cannot do both (Guyon’s conjecture; answered by the quadratic rough Heston model, by path-dependent volatility and by Guyon’s exact construction). A desk cares because VIX products and index products sensitive to volatility of volatility (cliquets, options on variance) are priced on exactly the dynamics a one-sided fit leaves untested.

What the interviewer is looking for: what each smile constrains, and the products at risk.

Interview question 12.6 ★★★ trader, risk

Two models fit today’s index surface equally well: one-factor Bergomi and rough Bergomi. Which of your products would they price differently, and how would you decide?

Solution

Solution of Interview question 12.6.

They agree on vanillas of the quoted expiries and differ on everything that depends on the volatility of forward variance at other horizons: short-dated options below the quoted expiries, forward-start options and cliquets (forward skew), options on variance and on the VIX. Decide by testing each against the instruments that measure those dynamics (weekly options, VIX futures and options, forward-start quotes) and by comparing the hedging P&L each model would have produced; hold a model reserve for the difference on the products that remain.

What the interviewer is looking for: the product list, a test against other markets, and a reserve.

Terms defined in this chapter

See all 2333 terms in the glossary