Quantitative Methods · Methods
19State-Space Models and the Kalman Filter
A desk hedges Brent crude with WTI and sets the hedge ratio by a 60-day rolling regression of Brent’s daily price changes on WTI’s. The ratio has not been stable: fitted year by year it went from 0.41 in 2014 to 0.94 in 2021 and 1.01 in 2026. When the relationship shifts, the rolling regression takes a month to cover half the move. A Kalman filter that treats the ratio as a hidden state drifting from day to day promises to follow shifts in days. Fitted by maximum likelihood to sixteen years of prices, the filter’s answer is sobering: the ratio drifts so slowly relative to the day’s noise that the optimal filter is itself a month-long average, and it hedges 0.7% better than the window. A faster filter is available, at a price the likelihood refuses to pay. This chapter is about state-space models and the filter that estimates their hidden states: the linear Gaussian model and the Kalman recursions, the likelihood they deliver, the smoother and the EM algorithm, particle filters for the models that are not linear or not Gaussian, and the time-varying hedge ratio that tests them.
19.1 Linear Gaussian state-space models
Definition 19.1 (State-space model, local level model)
A linear Gaussian state-space model has an observation equation , , and a state equation , , with independent noises and . The local level model is the scalar case : an observed series equals a random-walk level plus noise.
The form covers a great deal: ARMA models (chapter 17) with the lagged values as the state, a regression whose coefficients drift (, ), a term structure whose factors move over time, a fair value observed through noisy prints. For the hedge, is Brent’s daily price change, is WTI’s, and the state is the hedge ratio , a random walk: , .
Definition 19.2 (Signal-to-noise ratio)
In a local level model, or a random-walk coefficient model, the signal-to-noise ratio is , the variance of the state’s daily move over the variance of the observation noise.
19.2 The Kalman filter
Definition 19.3 (Kalman filter, innovation, Kalman gain)
The Kalman filter (Kalman, 1960) computes recursively the conditional law . Given , the innovation is , with variance ; the Kalman gain updates the state, and , and the prediction step gives and .
Proposition 19.4 (The Kalman filter is Gaussian conditioning)
If , then with the update above, and .
Proof. Given , is jointly normal with means , variances and , and covariance . The conditional law of a normal vector given one of its components has mean and variance . The prediction step applies the linear state equation with independent noise. ∎
It is the Bayesian update of chapter 14 run once a day, with yesterday’s posterior as today’s prior. A missing observation is handled by skipping the update: the prediction step alone widens . That is how the chapter treats 20 April 2020, when the EIA spot price of WTI settled at dollars: the WTI changes of that day and the next ( and dollars, against Brent’s and ) are not a hedge relationship but a storage crisis (commercial inventories at Cushing, Oklahoma, had risen by 27 million barrels since mid-March to 83% of working capacity, and the WTI futures contract traded below zero for the first time since trading began in 1983, according to the EIA), and both days are set to missing.
Proposition 19.5 (Steady-state gain of the local level model)
For the local level model with signal-to-noise ratio , the predicted variance converges to with , and the gain to . The filtered level is then an exponentially weighted average of past observations with weight on the newest.
Proof. In steady state , that is ; dividing by , . The recursion is an exponential average. ∎
For small the gain is about , so the filter’s memory is about days and its half-life (Figure 19.1). The signal-to-noise ratio, not the analyst, decides how fast the filter reacts. For a regression coefficient the same holds with multiplied by the regressor’s mean square.
Definition 19.6 (Prediction-error decomposition)
The prediction-error decomposition writes the likelihood of the observations as the product of the one-step predictive densities, , with the innovations and their variances from the filter.
The filter therefore delivers the likelihood of at the cost of one pass, and maximum likelihood fits the noise variances. On the hedge, one more refinement matters: daily oil price changes are far from homoskedastic (the standard deviation of WTI’s was 0.78 dollars in 2017 and 3.62 in 2026), and a constant lets the volatile years dictate the fit. The chapter takes with an exponentially weighted volatility of WTI’s changes known the day before, which is the same as filtering the scaled equation .
19.3 A time-varying hedge ratio
Fitted to the 4 097 daily changes from January 2010 to September 2026, the scaled model gives , and a signal-to-noise ratio of 0.00047; with the scaled regressor’s mean square of 1.12, the filter’s half-response to a shift is days, exactly the 30 days a 60-day window needs to cover half a shift. The filtered ratio’s uncertainty is itself large: its standard deviation is about 0.10, so the daily ratio is known only to (Figure 19.2).
Each estimate hedges day with the ratio known at the close of day . Over the 4 035 days after a 60-day start, the mean squared hedge error, in dollars squared per barrel per day, is 3.324 unhedged, 1.329 with the ratio of the first 60 days held fixed, 1.276 with the 60-day regression and 1.267 with the Kalman filter: a reduction of 0.7% against the window. Without the volatility scaling, the maximum-likelihood filter is 1.4% worse than the window, because the volatile years push its up tenfold. Making the filter faster is possible: multiplying by 10 cuts the half-response to 9.6 days and raises the hedge error by 0.6% relative to the window; multiplying it by 100 gives 3.0 days and 6.2% more error (Figure 19.3). In a simulation where the ratio jumps from 0.7 to 1.0 with the fitted noise levels, the likelihood’s filter and the window both need about fifty days to reach 0.85; the filter with 100 times the needs eight.
19.4 The smoother and expectation–maximisation
Definition 19.7 (Kalman smoother)
The Kalman smoother computes and its variance from the filter’s output by a backward pass (Rauch, Tung and Striebel, 1965): with .
The smoother answers a different question from the filter: not what the desk could have known, but what the ratio most likely was. It is the tool for research on past data (Figure 19.2 shows it smoother and earlier than the filter) and a source of look-ahead bias when it leaks into a backtest.
Definition 19.8 (Expectation–maximisation algorithm)
The expectation–maximisation algorithm (Dempster, Laird and Rubin, 1977) maximises a likelihood with unobserved variables by alternating an E-step, the expected complete-data log-likelihood given the data and the current parameters, and an M-step that maximises it. For state-space models the E-step is the smoother and the M-step for and has a closed form (Shumway and Stoffer, 1982).
Each EM step cannot decrease the likelihood: the complete-data log-likelihood’s expectation is a lower bound on the log-likelihood, up to a constant, that touches it at the current parameters, and the M-step raises the bound. EM is stable and needs no derivatives, but it can be slow. Started from and on the hedge model, it raises the log-likelihood from to in ten steps and to in forty, still 32 points short of the maximum that direct maximisation finds (Figure 19.4): the likelihood is flat along the direction that trades against .
19.5 Particle filters
Definition 19.9 (Particle filter, sequential importance resampling)
A particle filter represents the filtering law of a state by a weighted sample of particles. Sequential importance resampling, the bootstrap filter of Gordon, Salmond and Smith (1993), propagates each particle through the state equation, weights it by the density of the new observation given the particle, and resamples in proportion to the weights; the average weight at each step estimates the one-step predictive density, and their product the likelihood.
The particle filter needs neither linearity nor Gaussian noise, only the ability to simulate the state and evaluate the observation density; its price is Monte Carlo error and degeneracy, measured by the effective sample size of chapter 14 computed from the weights, . On a linear Gaussian local level of 500 observations, where the Kalman filter is exact, 2 000 particles give a log-likelihood of against the exact and filtered means within 0.07 of the exact ones (whose posterior standard deviation is 0.45), with an average effective sample size of 1 678.
The useful case is nonlinear. A stochastic-volatility model for daily EUR/USD returns lets the log variance follow an AR(1), , and observes : no Kalman filter applies. With 5 000 particles on the last 1 000 returns, a grid search of the particle likelihood (same random numbers at every point) picks and (log-likelihood , against at , : a flat surface). The filtered volatility follows the GARCH volatility of chapter 18 but reacts faster (Figure 19.5); the effective sample size averages 4 449 but falls to 23 on 3 April 2025, when the euro rose 2.7% after the US tariff announcement of 2 April (a risk-off episode in which, unusually, the dollar depreciated, as the ECB’s Financial Stability Review of November 2025 notes) and almost every particle was too calm to explain it.
19.6 Tutorial: the hedge ratio that moved
Goal. Track the Brent–WTI hedge ratio with a Kalman filter fitted by maximum likelihood, and measure what it buys against a rolling regression. End state: Figures 19.2 and 19.3, the signal-to-noise ratio 0.00047 and the 0.7% reduction in hedge error.
The filter, with missing observations and the prediction-error likelihood.
def kalman_filter(y, Z, T, H, Q, a0, P0) -> dict: """Univariate observations y_t (NaN = missing). Z: (m,) constant or (n, m) time-varying.""" y = np.asarray(y, dtype=float) T, Q = np.atleast_2d(np.asarray(T, float)), np.atleast_2d(np.asarray(Q, float)) m = T.shape[0] n = y.size a, P = np.asarray(a0, float).reshape(m).copy(), np.atleast_2d(np.asarray(P0, float)).copy() out = {k: np.zeros((n, m)) for k in ("a_pred", "a_filt", "K")} out.update({k: np.zeros((n, m, m)) for k in ("P_pred", "P_filt")}) v, F = np.full(n, np.nan), np.full(n, np.nan) ll = 0.0 for t in range(n): out["a_pred"][t], out["P_pred"][t] = a, P z = _Zt(Z, t, m) if not np.isnan(y[t]): v[t] = y[t] - z @ a F[t] = float(z @ P @ z) + float(H) K = P @ z / F[t] a = a + K * v[t] P = P - np.outer(K, z @ P) out["K"][t] = K ll += -0.5 * (math.log(2 * math.pi * F[t]) + v[t] ** 2 / F[t]) out["a_filt"][t], out["P_filt"][t] = a, P a = T @ a P = T @ P @ T.T + Q out.update({"v": v, "F": F, "loglik": ll}) return outListing 19.1. The Kalman filter with the prediction-error decomposition. code/firm/kalman/firm_kalman.py The streaming tracker the desk would run each evening.
class HedgeTracker: """Streaming time-varying regression y_t = beta_t x_t + e_t, beta_{t+1} = beta_t + u_t.""" def __init__(self, q: float, h: float, beta0: float = 1.0, p0: float = 1.0): self.q, self.h, self.beta, self.p = q, h, beta0, p0 def update(self, x: float, y: float) -> tuple[float, float]: used = self.beta if not (math.isnan(x) or math.isnan(y)): f = x * x * self.p + self.h k = self.p * x / f self.beta += k * (y - x * self.beta) self.p -= k * x * self.p self.p += self.q return used, self.betaListing 19.2. A streaming time-varying hedge ratio. code/firm/kalman/firm_kalman.py - Run
compare(),tradeoff(),em_path()andsv_grid()inqm_kalman.py, thenfig_kalman.py.
What to change next. Let the ratio mean-revert ( with a mean) instead of wandering, and compare likelihoods; add WTI’s lagged change as a second regressor to absorb the time difference between the two spot assessments.
19.7 Build: the Kalman toolkit
Purpose. Hidden quantities the miniature firm tracks every day (hedge ratios, fair values seen through noisy prints, drifting signal loadings) are filtered here, with their uncertainty.
Interface. kalman_filter(y, Z, T, H, Q, a0, P0); rts_smoother(kf, T); em(…); fit_mle(…); local_level_gain(q); particle_filter(y, n, init, step, loglik_obs, rng); HedgeTracker(q, h, beta0, p0).update(x, y).
Rules. Filters for trading use only data up to the previous close; smoothers are for research and never feed a backtest; missing data are NaN, not zeros; every particle filter reports its effective sample size.
Acceptance tests. code/firm/kalman/tests/: the likelihood and the smoothed states equal direct Gaussian conditioning on a short series; a missing observation skips the update; the steady-state gain formula; maximum likelihood and EM (monotone) recover a simulated local level; the particle filter matches the Kalman filter on a linear model; the streaming tracker reproduces the filter.
Stretch. Square-root filtering for numerical stability; the extended and unscented filters; particle marginal Metropolis–Hastings for the stochastic-volatility parameters.
Sources and further reading
- EIA, Today in Energy, on 2020 crude oil prices (negative WTI on 20 April 2020); ECB, Financial Stability Review, November 2025, special feature on the April 2025 tariff announcement.
- R. E. Kalman, “A new approach to linear filtering and prediction problems”, Journal of Basic Engineering 82, 1960.
- H. E. Rauch, F. Tung and C. T. Striebel, “Maximum likelihood estimates of linear dynamic systems”, AIAA Journal 3, 1965.
- A. P. Dempster, N. M. Laird and D. B. Rubin, “Maximum likelihood from incomplete data via the EM algorithm”, Journal of the Royal Statistical Society B 39, 1977.
- R. H. Shumway and D. S. Stoffer, “An approach to time series smoothing and forecasting using the EM algorithm”, Journal of Time Series Analysis 3, 1982.
- N. J. Gordon, D. J. Salmond and A. F. M. Smith, “Novel approach to nonlinear/non-Gaussian Bayesian state estimation”, IEE Proceedings F 140, 1993.
- US Energy Information Administration, WTI and Brent spot prices, via FRED (series DCOILWTICO and DCOILBRENTEU), accessed 24 September 2026.
19.8 Exercises
Exercise 19.1 ★
A local level model has and . What are the steady-state gain and the half-life of the filter’s memory?
Solution
Solution of Exercise 19.1.
: and . Old observations’ weights decay like , which halves in days.
Exercise 19.2 ★
Yesterday’s filtered hedge ratio is with variance ; , . Today WTI moves and Brent (scaled units). Update the ratio.
Solution
Solution of Exercise 19.2.
Predicted variance . Innovation , , gain . Updated ratio , variance .
Exercise 19.3 ★
Why does a missing observation increase the state’s variance, and by how much in the local level model?
Solution
Solution of Exercise 19.3.
Without an observation there is no update, only the prediction step, which adds the state noise: in the local level model the variance grows by that day, instead of first shrinking by the update. The uncertainty reflects that the state kept moving while nothing was seen.
Exercise 19.4 ★★
Show that for small the local level gain is approximately , and compare with an exponential moving average: which does the filter imply for ?
Solution
Solution of Exercise 19.4.
For small , and . The filtered level is , an EMA with ; for , and .
Exercise 19.5 ★★
Write the AR(1) plus noise, with , in state-space form and give its steady-state predicted variance equation.
Solution
Solution of Exercise 19.5.
State , , , , . The steady predicted variance solves .
Exercise 19.6 ★★
A particle filter’s normalised weights are 0.5, 0.3, 0.1, 0.05 and 0.05. What is the effective sample size?
Solution
Solution of Exercise 19.6.
particles’ worth out of five.
Exercise 19.7 ★★★
Coding. On the Brent–WTI data, scale the fitted by 0.1, 1, 10 and 100 and tabulate the hedge-error variance relative to the 60-day regression and the half-response time.
Solution
Solution of Exercise 19.7.
tradeoff(): : 0.987 of the window’s error, half-response 96 days; : 0.993, 30 days; : 1.006, 9.6 days; : 1.062, 3.0 days. Slower than the likelihood’s choice is marginally better here; faster costs noise.
Exercise 19.8 ★★★
Find the flaw. “We backtested a hedge using the Kalman-smoothed ratio and cut the hedge error by 15% against the rolling regression.”
Solution
Solution of Exercise 19.8.
The smoothed ratio for day uses prices after day , so the backtest hedges with information the desk could not have had. Use the filtered ratio from the previous close; the smoother belongs to research on what the relationship was, not to a simulation of trading.
19.9 Problem: The Hedge Ratio That Moved
Problem 19.1
Weekend problem — how fast should a hedge ratio move?
The data are the EIA’s daily WTI and Brent spot prices from FRED, 4 January 2010 to 22 September 2026 (4 097 daily changes). Brent’s daily price change is hedged with WTI’s; the hedge ratio for day must be known at the close of day .
Part I — The data.
- What happened on 20 April 2020, and how does the chapter treat it?
- How did the yearly regression ratio move between 2014, 2021 and 2026?
- Why scale the observation noise by WTI’s volatility?
- What are the unhedged and the fixed-ratio hedge-error variances?
- What does the 60-day regression achieve?
Part II — The filter.
- What , and signal-to-noise ratio does maximum likelihood give?
- What is the filter’s half-response to a shift, and how does it compare with the window’s?
- What hedge-error variance does the filter achieve, and what reduction against the window?
- What happens without the volatility scaling?
- How precisely is the daily ratio known?
Part III — Speed and noise.
- What do filters with multiplied by 10 and by 100 achieve?
- In the simulated jump from 0.7 to 1.0, how many days do the filters and the window need to reach 0.85?
- What does EM find from a poor start, and why is it slow?
- What are the filtered and rolling ratios at the end of the sample?
- When would a fast filter be worth its noise?
Part IV — Judgement.
- Should the desk replace its rolling regression?
- What else could improve the hedge more than the choice of filter?
- Why must the smoother stay out of the backtest?
- State the named result: the reduction in hedge-error variance and the estimated signal-to-noise ratio.
- In one sentence: what decides how fast a Kalman filter reacts?
Solution
Solution of Problem 19.1.
1. The EIA WTI spot price settled at dollars; WTI changed by and on 20 and 21 April while Brent changed by and . Both days are treated as missing: the filter skips the update, and the rolling regression and the evaluation skip the days. 2. 0.41 in 2014, 0.94 in 2021 and 1.01 in 2026. 3. WTI’s daily volatility ranged from 0.78 dollars (2017) to 3.62 (2026); with constant the volatile periods dominate the likelihood and push up; scaling makes the observation noise closer to homoskedastic. 4. 3.324 unhedged and 1.329 with the first 60 days’ ratio held fixed (dollars squared per barrel per day, 4 035 days). 5. 1.276. 6. , , signal-to-noise ratio 0.00047. 7. days, the same as the 60-day window’s 30. 8. 1.267, a reduction of 0.7%. 9. The fitted is ten times larger (0.0026 against in dollar units) and the filter hedges 1.4% worse than the window. 10. To about : the filtered ratio’s standard deviation is about 0.10. 11. : 0.6% more error than the window, half-response 9.6 days; : 6.2% more, 3.0 days. 12. About fifty days for the likelihood’s filter (50) and for the window (49); eight days with 100 times the (18 with ten times). 13. From , it reaches after ten steps and after forty, against the maximum : the likelihood is flat along a ridge trading against , where EM’s steps are short. 14. Filtered 1.20, rolling 1.28 (22 September 2026). 15. When shifts are large and rare compared with the day’s noise (a structural break), which a random walk with Gaussian steps does not describe; a model with occasional jumps in the ratio, or a change-point detector, would react fast only when needed. 16. Only for the uncertainty band, the principled tuning and the handling of missing days; not for the hedge error, which is almost the same. 17. Synchronising the prices (the two spot assessments are taken at different times), adding WTI’s lagged change, and hedging with futures rather than spot prices. 18. Because it uses future data: any backtest with it hedges with hindsight. 19. Named result: the hedge ratio that moved: on 2010–2026 data the maximum-likelihood Kalman hedge ratio reduces the hedge-error variance by 0.7% against the 60-day rolling regression, with an estimated signal-to-noise ratio of 0.00047, which makes the filter as slow as the window. 20. The ratio of the state’s noise to the observation noise, which the likelihood estimates from the data.
19.10 Interview questions
Interview question 19.1 ★ researcher, trader
Explain the Kalman filter to a trader in three sentences.
Solution
Solution of Interview question 19.1.
You believe something you cannot see, like a hedge ratio, drifts slowly; each day you predict it, see how wrong the prediction of today’s prices was, and move your belief toward the evidence by a fraction that depends on how noisy the prices are compared with how fast the hidden thing moves. You also track how uncertain your belief is. It is the optimal way to do this when everything is linear and Gaussian.
What the interviewer is looking for: Predict, compare, correct in proportion to relative noise, and carry the uncertainty.
Interview question 19.2 ★★ researcher
Derive the Kalman update for a scalar state from Gaussian conditioning.
Solution
Solution of Interview question 19.2.
Prior , observation , . are jointly normal with and . Conditioning: and ; the gain is .
What the interviewer is looking for: The joint-normal argument, not a memorised formula.
Interview question 19.3 ★★ researcher, trader
How do you choose the process noise of a Kalman filter tracking a hedge ratio?
Solution
Solution of Interview question 19.3.
Estimate it with the observation noise by maximum likelihood from the prediction-error decomposition, then check the hedge out of sample and the speed-noise trade-off; choosing it by hand to “react in days” buys noise. Scale the observation noise with volatility first, or the fit is dominated by volatile periods.
What the interviewer is looking for: Likelihood-based tuning, the trade-off, and heteroskedasticity.
Interview question 19.4 ★★ researcher, mle
What is the difference between filtering and smoothing, and why does it matter in backtests?
Solution
Solution of Interview question 19.4.
Filtering conditions on data up to ; smoothing on all data, past and future. The smoothed path is better for describing history and wrong for trading simulations, where it is look-ahead bias.
What the interviewer is looking for: Information sets and look-ahead bias.
Interview question 19.5 ★★ mle, researcher
When would you use a particle filter instead of a Kalman filter, and what can go wrong?
Solution
Solution of Interview question 19.5.
When the model is nonlinear or non-Gaussian: stochastic volatility, jumps, heavy-tailed noise, discrete regimes. It can degenerate (few particles carry the weight) on surprising observations, its likelihood is noisy (which complicates parameter estimation), and its cost grows with the state dimension.
What the interviewer is looking for: Nonlinearity as the reason, degeneracy and Monte Carlo noise as the risks, effective sample size as the monitor.
Interview question 19.6 ★★★ researcher
Why does each EM iteration not decrease the likelihood?
Solution
Solution of Interview question 19.6.
For any law of the hidden variables, . The E-step sets to the posterior at , making the KL term zero, so the bound touches the likelihood there; the M-step raises the bound; the likelihood at is at least the raised bound, hence at least the likelihood at .
What the interviewer is looking for: The lower bound and the KL gap.