Quantitative Finance · Book 15 · Technology

Research, Data and Risk Platforms

Research, Data and Risk Platforms · Technology

19Pricing-Library Architecture

A client asked on Friday for a price on a worst-of autocallable with memory coupons and a lookback on the knock-in: the capital is at risk if the worst performance was ever below 60% on an observation date, not only at maturity. The library had no class for it. By Monday it was priced from eleven lines of payoff script, on the same market objects, model and Monte Carlo engine as every other trade in the library, with Greeks, scenarios and batches for free. This chapter is about the architecture that makes that possible: a library in layers, whose engines do not know which products they price, and a small language in which a product is data.

19.1 The analytics library and its layers

Definition 19.1 (Analytics library)

An analytics library is the firm’s shared code for valuing and risk-managing instruments: market objects, models, numerical engines and instruments, separated so that each can be replaced or extended without the others, and called the same way by pricing tools, risk systems and research.

Book 5 built one (firm.pricing, chapter 28) and froze its interface for the rest of the series: instruments are descriptions of cash flows; market data is an immutable snapshot; a model says how the market evolves; an engine computes a value for an instrument under a model on a snapshot; and every Greek, scenario and batch is one repricing function applied to bumped snapshots. The layers are the architecture: a new engine prices old instruments, a new instrument uses old engines, and a risk system (chapter 18, and Book 6’s engine) needs to know nothing but the repricing function.

The library’s layers and the script engine’s place in them. A scripted instrument is an instrument; the script engine is an engine, registered for the Black–Scholes model; everything above — Greeks, scenarios, batches — applies unchanged.
Figure 19.1. The library’s layers and the script engine’s place in them. A scripted instrument is an instrument; the script engine is an engine, registered for the Black–Scholes model; everything above — Greeks, scenarios, batches — applies unchanged.

19.2 Instruments, market objects, models and engines

Definition 19.2 (Market object)

A market object is a library object that represents one piece of market state — a quote, a discount curve, a volatility surface, a correlation — and answers questions about it (a discount factor, an implied volatility), so that engines read the market through a fixed interface whatever its construction.

In firm.pricing a snapshot holds the market objects and a bump makes a new snapshot with one object replaced; the Monte Carlo engine reads forwards, variances and correlations from it and generates paths from a seed fixed per instrument (common random numbers, Book 5), so that a bumped price differs from the base price by the bump alone. The script engine of this chapter reuses those paths: it is the Monte Carlo engine with the payoff taken from a script instead of from a class.

19.3 Payoff scripting

Definition 19.3 (Payoff scripting language, payoff script)

A payoff scripting language is a small programming language in which a product’s cash flows are written as events on observation dates — conditions on the path, payments, state carried from date to date — and evaluated by a generic engine on simulated paths. A payoff script is one product written in it: data that the library prices, not code that it must be extended with.

The chapter’s language (firm.payoffdsl) has what the chapter’s products need and nothing more: state variables, blocks run at the dates of named schedules, assignments, if/else, pay and stop, arithmetic and comparisons, and min, max and abs. At each date a block sees S (the first underlying’s spot), perf (the worst performance) and t. A tokenizer turns the text into tokens, a recursive-descent parser turns them into a syntax tree with an error position for anything it cannot read, and the schedule is extracted from the tree: every block’s dates, in date order, in script order on the same date.

From a script to a price: the tokenizer and the parser build the tree, the schedule is read from its blocks, and either evaluator runs the tree on the library’s paths for the schedule’s dates.
Figure 19.2. From a script to a price: the tokenizer and the parser build the tree, the schedule is read from its blocks, and either evaluator runs the tree on the library’s paths for the schedule’s dates.
AUTOCALL = """state missed = 0;
at all {
  if perf >= barrier { pay cpn + missed; missed = 0; }
  else { missed = missed + cpn; }
}
at calls { if perf >= trigger { pay 100; stop; } }
at last {
  if perf < protection { pay 100 * min(perf, 1); } else { pay 100; }
}"""
Listing 19.1. The worst-of Phoenix autocallable with memory coupons as a payoff script: coupons on every observation date, an autocall on all but the last, the capital at risk at maturity. code/platforms/19-pricing-library-architecture/python/pl_payoffdsl.py
    def stmt(self):
        t = self.peek()
        if t.text == "if":
            self.take("if")
            cond, then = self.expr(), self.block()
            other = ()
            if self.peek("else"):
                self.take("else")
                other = (self.stmt(),) if self.peek("if") else self.block()
            return If(cond, then, other)
        if t.text == "pay":
            self.take("pay")
            e = self.expr()
            self.take(";")
            return Pay(e)
        if t.text == "stop":
            self.take("stop")
            self.take(";")
            return Stop()
        name = self.take(kind="name").text
        self.take("=")
        e = self.expr()
        self.take(";")
        return Assign(name, e)
Listing 19.2. The parser’s statements: if/else, pay, stop and assignment; every token keeps its line and column for error messages. code/firm/payoffdsl/firm_payoffdsl.py

The client’s product changes three lines: a state variable lowest, updated with lowest = min(lowest, perf) at every observation, and the maturity test on lowest instead of perf. It prices at 94.25 per 100, 1.63 below the plain autocallable’s 95.88: the lookback knock-in triggers on paths that dip below 60% and recover.

19.4 Two ways to evaluate a script

The simplest evaluator walks the syntax tree once per path and per event, with scalar values: an interpreter. It is easy to trust and slow. The compiled evaluator walks the tree once per event with arrays over all paths: every statement runs under a mask of the paths it applies to, an if splits the mask, a pay adds under it, a stop removes paths from the living set (Listing 19.3). The two give the same numbers to the last digit (the tests compare them); the compiled one is 19 to 34 times faster (Figure 19.3).

def compile_program(program: Program):
    """The same semantics on all paths at once: each statement under a mask of its paths."""

    def block(stmts, env, mask, pay, alive):
        for s in stmts:
            live = mask & alive
            if isinstance(s, Assign):
                old = env[s.name] if s.name in env else 0.0
                env[s.name] = np.where(live, _eval(s.expr, env), old)
            elif isinstance(s, Pay):
                pay += np.where(live, _eval(s.expr, env), 0.0)
            elif isinstance(s, Stop):
                alive &= ~live
            else:
                c = np.asarray(_eval(s.cond, env), bool)
                block(s.then, env, live & c, pay, alive)
                block(s.other, env, live & ~c, pay, alive)

    def run(perf, spot, events, times, dfs, params) -> np.ndarray:
        n = perf.shape[0]
        env = {k: np.full(n, float(v)) for k, v in params.items()}
        for name, e in program.states:
            env[name] = np.broadcast_to(np.asarray(_eval(e, env), float), (n,)).copy()
        pv, alive = np.zeros(n), np.ones(n, bool)
        for j, k in events:
            env.update(S=spot[:, j], perf=perf[:, j], t=np.full(n, times[j]))
            pay = np.zeros(n)
            block(program.events[k][1], env, np.ones(n, bool), pay, alive)
            pv += pay * dfs[j]
        return pv

    return run
Listing 19.3. The compiled evaluation: masks instead of branches, arrays instead of loops. code/firm/payoffdsl/firm_payoffdsl.py
Time to price the autocallable script by the interpreter and by the compiled evaluator, path generation included, best of three, one thread. The compiled version is 18.8 times faster at 1 000 paths, 34.3 times at 10 000 and 33.4 times at 30 000. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2, machine otherwise idle. Data: bench_payoffdsl.py.
Figure 19.3. Time to price the autocallable script by the interpreter and by the compiled evaluator, path generation included, best of three, one thread. The compiled version is 18.8 times faster at 1 000 paths, 34.3 times at 10 000 and 33.4 times at 30 000. Measured on a laptop (Intel Core Ultra 7 155H) under WSL2, machine otherwise idle. Data: bench_payoffdsl.py.

The script engine is registered in the library like any engine (Listing 19.4): it parses the script, extracts the schedule, asks the Monte Carlo engine for paths on the schedule’s dates with the instrument’s seed, and evaluates. Because the seed depends only on the instrument’s identifier, a script with the same identifier as a library instrument sees the same paths.

class ScriptEngine:
    n_paths: int = 100_000
    mode: str = "compiled"
    name: str = "Script"

    def supports(self, inst, model) -> bool:
        return isinstance(inst, ScriptedInstrument) and type(model) is FP.BlackScholes

    def cashflows(self, inst: ScriptedInstrument, model, md: FP.MarketData) -> np.ndarray:
        prog = parse(inst.script)
        ev = schedule(prog, dict(inst.dates))
        dates = sorted({d for d, _k in ev})
        idx = {d: j for j, d in enumerate(dates)}
        names = inst.underlyings or (inst.underlying,)
        refs = [md.spots[u] for u in names]
        mc = FP.MonteCarloEngine(n_paths=self.n_paths)
        s = mc.paths(inst, md, list(names), dates, refs, model)
        init = np.array(inst.initial or refs)
        perf = (s / init[None, None, :]).min(axis=2)
        times = np.array([md.t(d) for d in dates])
        dfs = np.array([md.df(inst.curve(md), d) for d in dates])
        events = [(idx[d], k) for d, k in ev]
        run = compile_program(prog) if self.mode == "compiled" else Interpreter(prog).run
        return inst.notional * run(perf, s[:, :, 0], events, times, dfs, dict(inst.params))

    def price(self, inst, model, md) -> float:
        return float(self.cashflows(inst, model, md).mean())
Listing 19.4. The script engine: schedule, the library’s paths, discounting on the instrument’s curve, and the evaluator. code/firm/payoffdsl/firm_payoffdsl.py
productscripts.e.librarydifference (s.e.)library engine
European call11.39530.057011.34850.82analytic
Asian call (12 monthly fixings)6.80120.03276.80120.00Monte Carlo
daily down-and-out call11.03730.056811.02260.26closed form (BGK)
worst-of autocallable95.87500.071795.87500.00Monte Carlo
autocallable, lookback knock-in94.24740.0726——none
Table 19.1. Scripts against the library’s own engines, 100 000 paths. Against an analytic or closed-form engine the difference is Monte Carlo error, under one standard error; against the library’s Monte Carlo engine on common random numbers it is zero, because the paths and the payoff arithmetic are the same.

Example 19.4 (Priced by Monday)

The scripted autocallable prices at 95.8750 per 100 and the library’s own autocallable engine at 95.8750: the difference is 0.00 standard errors, because both evaluate the same payoff on the same paths. Their deltas to the first underlying are equal too, 0.1518 per unit of spot per 100 of notional, since the library’s bump-and-reprice Greeks work on any engine. The client’s lookback variant, which the library cannot express, prices at 94.2474 (±0.0726\pm0.0726) with the same machinery. The compiled evaluator makes the script engine as fast as a hand-written array engine; the interpreter, 22 to 37 times slower, remains the reference that checks it.

19.5 Lazy recalculation and observers

Definition 19.5 (Observer pattern)

The observer pattern lets an object (the observable) keep a list of dependants (observers) and notify them when it changes, without knowing what they are; a curve observes its quotes, a price observes its curve.

Definition 19.6 (Lazy recalculation)

Lazy recalculation marks a result stale when a notification arrives and recomputes it only when the result is next asked for, so that a burst of market changes costs one recalculation per reader instead of one per change.

QuantLib’s LazyObject is the best-known implementation: a “framework for calculation on demand and result caching” that is both an observer and an observable. The chapter’s LazyPrice is the same idea in twenty lines (Listing 19.5): a European call observing a rate quote is computed once (11.35), the quote changes three times, and the next read recomputes once (11.59) — two calculations for four states of the market. The snapshot design of firm.pricing is the other answer to the same problem: nothing is mutable, so nothing is stale; a new snapshot is a new question. Libraries serving interactive tools favour observers; libraries serving batch risk favour snapshots.

class LazyPrice(Observable):
    """A price marked stale when an input changes, recomputed only when it is asked for."""

    def __init__(self, compute, *inputs):
        super().__init__()
        self.compute, self.inputs = compute, inputs
        self.cached, self.stale, self.calculations = None, True, 0
        for q in inputs:
            q.register(self)

    def update(self) -> None:
        if not self.stale:
            self.stale = True
            self.notify()                   # observers of this price are stale too

    @property
    def value(self) -> float:
        if self.stale:
            self.cached = self.compute(*(q.value for q in self.inputs))
            self.calculations += 1
            self.stale = False
        return self.cached
Listing 19.5. A lazy price: stale on notification, recomputed on demand, and itself observable by what depends on it. code/firm/payoffdsl/firm_payoffdsl.py

19.6 Bindings to other languages

A library written once is called from many places: Python research, a C++ pricing service, a spreadsheet on a trader’s desk. Chapter 9 measured what a binding costs per call; for a library the lesson is to bind at the level of whole requests — a batch of trades, a snapshot — never at the level of an operation. Book 5’s library has C++20 and Rust twins of its core, tested against the Python reference on the same inputs; a payoff script is portable by construction, since it is text that any implementation of the language can price, which is why scripts travel between front-office tools, risk systems and the canonical trade model of chapter 21 more easily than classes do.

19.7 Tutorial: a language, an engine, five products

Goal. Write products as scripts, price them through the library, check them against the library’s engines, and measure the two evaluators. End state: Table 19.1 and Figure 19.3.

  1. Parse: firm_payoffdsl.parse on each script; break one and read the error position.
  2. Price: firm.pricing.price on a ScriptedInstrument; greeks on it.
  3. Validate: pl_payoffdsl.validation().
  4. Measure: bench_payoffdsl.py.
  5. Observe: lazy_demo().

What to change next. Add a worst function over named underlyings and a daily lookback between observation dates; compile a script to Python source instead of closures and measure again.

19.8 Build: the payoff language

Purpose. Products the library has no class for, priced, bumped and batched like those it has, from scripts that can be stored, reviewed and moved between systems.

Interface. parse, schedule, Interpreter, compile_program, ScriptedInstrument, ScriptEngine (registered for Black–Scholes), ScriptError; Quote, LazyPrice.

Rules. The library’s paths and seeds, never the engine’s own; discounting on the instrument’s curve; the interpreter and the compiled evaluator agree on every path; errors name a line and a column.

Acceptance tests. code/firm/payoffdsl/tests/: parsing and error positions; the event schedule’s order; the interpreter equal to the compiled evaluator on a path-dependent script; a scripted European equal to the library’s Monte Carlo price and close to its analytic one, with its delta; lazy prices and chained observers.

Stretch. Several underlyings by name; adjoint (AAD) Greeks through the compiled script (Book 4, chapter 28); a script library with versioned, reviewed products.

Sources and further reading

  • QuantLib reference documentation, LazyObject.
  • One Quant Book 5, chapters 4, 18 and 28 (Greeks, autocallables, the library); One Quant Book 4, chapter 26 (Monte Carlo).

19.9 Exercises

Exercise 19.1 ★

Write a digital call paying 10 at expiry if the spot is above 105 as a payoff script.

Solution

Solution of Exercise 19.1.

at expiry { if S > 105 { pay 10; } } with the schedule expiry holding the expiry date. On the chapter’s market it prices at 4.09, equal to the library’s digital option priced by its Monte Carlo engine on the same paths.

Exercise 19.2 ★

Why is the scripted autocallable’s price exactly the library’s, while the scripted European differs from the analytic price by 0.82 standard errors?

Solution

Solution of Exercise 19.2.

The autocallable is compared with the library’s Monte Carlo engine on the same paths (same identifier, same seed, same dates) and the same payoff arithmetic: the numbers are identical. The European is compared with the analytic formula, which has no Monte Carlo error; the script’s estimate differs from it by its own sampling error, 0.82 standard errors here.

Exercise 19.3 ★

In the autocallable script, why must the coupon block come before the autocall block on the same date?

Solution

Solution of Exercise 19.3.

Because the term sheet pays the coupon on an observation date and then redeems if the note is called: with the autocall first, stop would end the path before the coupon of the call date was paid, and the script would undervalue every called path by a coupon (and its memory).

Exercise 19.4 ★★

How does the compiled evaluator handle stop, and what would go wrong if it simply skipped the rest of the block?

Solution

Solution of Exercise 19.4.

It removes the paths under the current mask from the living set; every later statement and event runs under mask & alive, so those paths pay and change nothing more. Skipping the rest of the block would skip it for all paths, including those for which the stop did not apply.

Exercise 19.5 ★★

Why does the speed-up of the compiled evaluator grow with the number of paths and then level off?

Solution

Solution of Exercise 19.5.

The compiled evaluator pays a fixed cost per event (walking the tree once) and a small cost per path; the interpreter pays the walk per path. At few paths the fixed cost and path generation (common to both) weigh more; as paths grow the ratio approaches the ratio of per-path costs, and levels off once both are dominated by array work and memory traffic.

Exercise 19.6 ★★

Observers or snapshots: which would you choose for a trader’s pricing screen, and which for the overnight risk run?

Solution

Solution of Exercise 19.6.

Observers for the screen, where one market quote changes at a time and only the prices on screen need recomputing; snapshots for the risk run, where thousands of bumped markets are evaluated independently and in parallel, and nothing should depend on the order of updates.

Exercise 19.7 ★★★

Coding. Price the lookback autocallable with the knock-in level at 50% and at 70%. How does the price move, and what does that say about the product’s sensitivity to the barrier?

Solution

Solution of Exercise 19.7.

97.64 at 50%, 94.25 at 60%, 92.28 at 70%: about 5.4 points over twenty points of barrier. A lookback knock-in is very sensitive to the barrier, since any observation below it counts; the barrier level is a pricing parameter to negotiate and a risk to hedge.

Exercise 19.8 ★★★

Find the flaw. “Our scripting engine draws its own random numbers, so script prices are independent of the library and a good cross-check.”

Solution

Solution of Exercise 19.8.

Independent random numbers make the script and the library differ by Monte Carlo noise, which hides small payoff errors; and they give up common random numbers, so bumped script prices are noisy and Greeks unstable. Draw the library’s paths with the instrument’s seed; cross-check by comparing with engines on the same paths (exact) and with closed forms (within the standard error).

19.10 Problem: Priced by Monday

Problem 19.1

Weekend problem — a product the library has no class for

firm.pricing, firm.payoffdsl and the chapter’s five products.

Part I — The library.

  1. What are the library’s four layers, and what connects them?
  2. What does a market object answer, and why an immutable snapshot?
  3. Why do bumped prices differ by the bump alone?
  4. What must an engine provide to be registered?
  5. Why does a risk system need nothing but the repricing function?

Part II — The language.

  1. What can a script say, and what does each block see?
  2. How is the event schedule extracted?
  3. How do the interpreter and the compiled evaluator differ?
  4. How does the client’s product differ from the plain autocallable in the script?
  5. What happens to a syntax error?

Part III — The numbers.

  1. Give the script and library prices of the five products and their differences in standard errors.
  2. Why are two differences exactly zero?
  3. What are the autocallable’s deltas by script and by library?
  4. What is the compiled evaluator’s speed-up?
  5. What does the lazy price’s demonstration count?

Part IV — The verdict.

  1. State the named result: the scripted autocallable’s price and delta against the library’s autocallable engine on common random numbers, their difference in standard errors, and the speed-up of the compiled script over the interpreted one.
  2. What would you require before a scripted product is traded?
  3. When should a product get its own class and engine instead of a script?
  4. How would you store and version scripts?
  5. In one sentence: what does a payoff language buy a pricing library?
Solution

Solution of Problem 19.1.

  1. Instruments, market objects, models and engines; the repricing function on snapshots connects them.
  2. Questions about one piece of market state (a discount factor, a volatility); a snapshot never changes, so a bump makes a new one and nothing is stale.
  3. Common random numbers: the seed depends only on the instrument, so base and bumped prices use the same paths.
  4. supports and price for an (instrument, model) pair.
  5. Because every Greek, scenario and batch is that function applied to bumped snapshots.
  6. State, events on schedules, assignments, conditions, payments and stops; S, perf, t, the parameters and the state.
  7. From the at blocks: every date of every block, sorted by date, in script order within a date.
  8. The interpreter walks the tree per path with scalars; the compiled evaluator walks it per event with arrays and masks.
  9. A state variable for the lowest performance, updated at every observation, and the maturity test on it.
  10. The parser raises an error with its line and column.
  11. European 11.3953 against 11.3485 (0.82); Asian 6.8012 against 6.8012 (0.00); down-and-out 11.0373 against 11.0226 (0.26); autocallable 95.8750 against 95.8750 (0.00); lookback 94.2474, no library class.
  12. Both are compared with the library’s Monte Carlo engine on the same paths and payoffs.
  13. 0.1518 per unit of spot per 100 of notional, both.
  14. 18.8 times at 1 000 paths, 33.4 at 30 000, 34.3 at most (at 10 000).
  15. Two calculations for four states of the market: one at the first read, one at the read after three changes.
  16. Named result. The scripted autocallable prices at 95.8750 per 100 with a delta of 0.1518, exactly as the library’s autocallable engine on common random numbers (0.00 standard errors apart); the compiled script runs 22 to 37 times faster than the interpreted one.
  17. Validation against library engines and closed forms, review of the script, its storage with a version, and model approval.
  18. When it needs a numerical method the script engine does not offer (early exercise by regression, a PDE), or when it trades in volume and deserves a dedicated, faster engine.
  19. In version control as text, with the trade referring to a script version and its parameters; every change reviewed and repriced.
  20. New products as data, priced and risk-managed by the machinery the library already trusts.

19.11 Interview questions

Interview question 19.1 ★ developer

Why separate instruments, models and engines in a pricing library?

Solution

Solution of Interview question 19.1.

So that each varies independently: a new model prices old products, a new engine speeds up old models, a new product uses existing engines; risk and scenarios work on any combination through one repricing function.

What the interviewer is looking for: Separation of concerns, one repricing path.

Interview question 19.2 ★★ developer, researcher

How would you price a new exotic by Monday without writing a new engine?

Solution

Solution of Interview question 19.2.

Write it as a payoff script on the library’s scripting engine, validate against the closest products the library prices, check Greeks and limits, and have it reviewed; a new engine can come later if volume justifies it.

What the interviewer is looking for: Scripting and validation.

Interview question 19.3 ★★ developer

How do you make a Monte Carlo payoff evaluator fast in Python?

Solution

Solution of Interview question 19.3.

Evaluate over all paths at once with arrays: conditions become masks, payments masked sums, termination a living set; generate paths once; avoid per-path Python loops, and fall back to compiled code if the tree walk itself becomes the cost.

What the interviewer is looking for: Vectorisation by masks.

Interview question 19.4 ★★ developer

What is lazy evaluation in a pricing library, and when is it a bad idea?

Solution

Solution of Interview question 19.4.

Results marked stale on input changes and recomputed when read; bad when reads are constant (no saving), in parallel batch risk (shared mutable state), and when notification chains grow long and their order matters.

What the interviewer is looking for: Observers against immutable snapshots.

Interview question 19.5 ★★★ developer

Design a payoff scripting language for a derivatives desk.

Solution

Solution of Interview question 19.5.

A small grammar (events on schedules, state, conditions, payments, stops), a parser with error positions, schedule extraction, an evaluator on the library’s paths with common random numbers, registration as an engine, a validation suite against library products, versioned storage and review of scripts.

What the interviewer is looking for: Small language, library paths, validation, versioning.

Terms defined in this chapter

See all 2333 terms in the glossary