---
title: "Culture and Decision-Making"
book: "The Desk and the Firm"
subject: quant
language: en
chapter: 26
exercises: 8
source: https://one-course.com/books/quant/16/en/chapter/26-culture-and-decision-making
---

# Chapter 26 — Culture and Decision-Making

In a two-year geopolitical forecasting tournament in which five university-based research groups competed to produce the most accurate probabilities, one group reported that three interventions made its forecasters better: training them in probability, letting them share information and argue their reasons in teams, and putting the best of them together. Each improved both how well the forecasters’ probabilities matched outcomes and how sharply they separated events that happened from those that did not; combined with statistical aggregation, the approach produced the best forecasts two years in a row. Writing a number down, being scored on it and arguing in the open is most of what a betting culture is, and a trading firm is a place where every decision is a bet.

## 26.1 Betting cultures: probabilities, not opinions

A trading desk decides under uncertainty all day: whether a signal is real, whether a strategy is broken (chapter 8), whether to hire, to enter a market (chapter 24), to raise a limit (chapter 12). A culture that states these beliefs as probabilities can score them later; one that states them as opinions (“I think it’ll work”) cannot, and learns only from outcomes, which mix skill and luck (chapter 10).

**Definition 26.1 (Decision journal).**

A *decision journal* records each significant decision when it is made: the options considered, the probabilities and expected values assigned to their outcomes, the reasons, who decided, and the date; the outcome is added later, so that decisions can be scored on the information available at the time.

## 26.2 Scoring forecasts and decisions

**Definition 26.2 (Brier score).**

The *Brier score* of probability forecasts $p_i$ of binary outcomes $y_i\in\{0,1\}$ is the mean squared error $\frac1N\sum_i(p_i-y_i)^2$: zero for perfect forecasts, 0.25 for a constant forecast of one half, lower is better.

Murphy’s decomposition (1973) splits the score into three parts: the uncertainty of the events themselves, $\bar o(1-\bar o)$, which no forecaster controls; the reliability, the mean squared gap between the forecasts in each probability bin and the frequency with which those events happened, which calibration training reduces; and the resolution, the spread of the bins’ observed frequencies around the overall one, which knowledge increases. The score is reliability minus resolution plus uncertainty ([Listing 26.1](#lst-fm-culture-and-decision-making-score)); Book 7’s calibration curve draws the reliability term.

```python
def brier(p, y):
    p, y = np.asarray(p, float), np.asarray(y, float)
    return float(np.mean((p - y) ** 2))


def _bins(p, bins):
    return np.clip((np.asarray(p, float) * bins).astype(int), 0, bins - 1)


def calibration(p, y, bins=10):
    p, y = np.asarray(p, float), np.asarray(y, float)
    b = _bins(p, bins)
    out = []
    for k in range(bins):
        m = b == k
        if m.any():
            out.append((float(p[m].mean()), float(y[m].mean()), int(m.sum())))
    return out


def murphy(p, y, bins=10):
    p, y = np.asarray(p, float), np.asarray(y, float)
    n, o = len(p), float(y.mean())
    rel = res = 0.0
    for f, ok, nk in calibration(p, y, bins):
        rel += nk * (f - ok) ** 2 / n
        res += nk * (ok - o) ** 2 / n
    return {"brier": brier(p, y), "reliability": rel, "resolution": res, "uncertainty": o * (1 - o)}
```

***Listing 26.1.** The Brier score, the calibration bins and Murphy’s decomposition. code/firm/decisionlog/firm_decisionlog.py*

**Definition 26.3 (Forecast aggregation).**

*Forecast aggregation* combines several forecasters’ probabilities for the same event into one: by averaging the probabilities, by averaging their log-odds, or by averaging the log-odds and multiplying them by a factor above one (extremising) to undo the moderation that averaging produces when the forecasters hold different information.

**Proposition 26.4 (Five opinions or one).**

If $K$ forecasters’ errors in log-odds each have variance $\sigma^2$ and pairwise correlation $\rho$, the error of their average has variance

$$
\sigma^2\Bigl(\rho+\frac{1-\rho}{K}\Bigr),
$$

which falls from $\sigma^2$ for one forecaster to $\sigma^2/K$ for independent ones, and stays at $\sigma^2$ whatever $K$ when $\rho=1$.

**Proof.** $\mathrm{Var}\bigl(\frac1K\sum_ke_k\bigr)=\frac1{K^2}\bigl(K\sigma^2+K(K-1)\rho\sigma^2\bigr)=\sigma^2\bigl(\rho+\frac{1-\rho}K\bigr)$. ∎

A discussion before forecasting correlates the errors: people converge on the loudest argument, the first number spoken, or the senior person’s view. The value of five opinions is in their independence, which is why good forecasting practice collects the numbers before the discussion and the discussion before a second round.

## 26.3 Tutorial: five opinions or one

**Goal.** Score a year of a desk’s forecasts, decompose the scores, and compare aggregating five independent forecasts with recording one consensus after a discussion. **End state:** calibration curves ([Figure 26.1](#fig-fm-culture-and-decision-making-calib)) and the [Brier scores](#def-fm-culture-and-decision-making-brier) by rule ([Figure 26.3](#fig-fm-culture-and-decision-making-rho)).

1. **The events.** 400 events a year, with true probabilities drawn from a Beta(2, 2); the truth itself scores 0.214, and a constant one half 0.25 ( `fm_culture.year` , synthetic).
2. **The forecasters.** Five see the truth’s log-odds with noise of standard deviation 0.8, and stretch their log-odds by 1.0 to 1.6: all but the first are overconfident.
3. **The scores.** `firm.decisionlog.brier` , `murphy` and `calibration` ; `aggregate` three ways.
4. **The discussion.** `fm_culture.five_or_one` correlates the forecasters’ errors at $\rho$ from 0 to 1 and scores the consensus over twenty simulated years.

![Calibration curves for a year of 400 forecasts: forecaster 5, who stretches its log-odds by 1.6, gives high probabilities to events that happen far less often; the average of five forecasters is close to the diagonal. Data: firm.decisionlog.calibration.](https://one-course.com/images/onecourse/chapters/quant-16/fm-culture-and-decision-making/fig-0c45eaa676a5.svg)

***Figure 26.1.** Calibration curves for a year of 400 forecasts: forecaster 5, who stretches its log-odds by 1.6, gives high probabilities to events that happen far less often; the average of five forecasters is close to the diagonal. Data: `firm.decisionlog.calibration`.*

In the chapter’s year the five forecasters score between 0.225 and 0.262, against 0.214 for the truth itself. The most overconfident, forecaster 5, has a reliability term of 0.040 against 0.017 for forecaster 1, with about the same resolution (0.026 and 0.027): its extra error is all miscalibration. Averaging the five probabilities scores 0.220 and averaging their log-odds 0.223, both better than every forecaster ([Figure 26.2](#fig-fm-culture-and-decision-making-brier)); the average’s reliability term falls to 0.010 and its resolution rises to 0.039, since the independent errors partly cancel. Extremising the average makes it worse (0.237): these forecasters are already overconfident, and extremising only helps forecasters whose average is too timid.

![Brier scores of a year’s 400 forecasts: five forecasters, three ways of aggregating them, and the truth’s own probabilities, the floor no forecaster can beat on average. Data: fm_culture.scores.](https://one-course.com/images/onecourse/chapters/quant-16/fm-culture-and-decision-making/fig-ef927973e67e.svg)

***Figure 26.2.** [Brier scores](#def-fm-culture-and-decision-making-brier) of a year’s 400 forecasts: five forecasters, three ways of aggregating them, and the truth’s own probabilities, the floor no forecaster can beat on average. Data: `fm_culture.scores`.*

![The Brier score of the desk’s forecast, averaged over twenty simulated years: five independent forecasts aggregated, against one consensus whose errors a discussion has correlated at . Data: fm_culture.five_or_one.](https://one-course.com/images/onecourse/chapters/quant-16/fm-culture-and-decision-making/fig-bbdfbb3ce56f.svg)

***Figure 26.3.** The [Brier score](#def-fm-culture-and-decision-making-brier) of the desk’s forecast, averaged over twenty simulated years: five independent forecasts aggregated, against one consensus whose errors a discussion has correlated at $\rho$. Data: `fm_culture.five_or_one`.*

Over twenty simulated years five independent forecasts average a [Brier score](#def-fm-culture-and-decision-making-brier) of 0.207; a consensus after a discussion that correlates the errors at 0.5 scores 0.224, and at 1 (everyone repeats the same view) 0.237 ([Proposition 26.4](#prop-fm-culture-and-decision-making-rho)). The difference, 0.017 at $\rho=0.5$, is more than half the gap between the average forecaster’s score in the chapter’s year and the truth’s: the discussion throws away most of what the independent forecasts were worth.

## 26.4 Post-mortems and pre-mortems

**Definition 26.5 (Outcome bias).**

*Outcome bias* is judging the quality of a decision by its outcome rather than by the information and reasoning available when it was made, so that good decisions that lost are called mistakes and bad decisions that won are rewarded.

Baron and Hershey (1988) studied the bias experimentally; the [decision journal](#def-fm-culture-and-decision-making-journal) is its antidote. In the chapter’s year forecaster 2, the best, bets at even odds on the 362 events it rates above 55% or below 45%: every bet has a positive expected value by its own forecast, and 123 of them, 34%, lose. A reviewer who judges by outcome calls a third of the desk’s good decisions bad. The same journal shows forecaster 2’s real flaw: it wins 66% of its bets while its forecasts average 77.7%, overconfidence that only a record of stated probabilities can reveal.

**Definition 26.6 (Pre-mortem, red team).**

A *pre-mortem* is a review held before a decision is implemented, in which the team imagines that it has failed and writes down why, to surface risks that optimism hides. A *red team* is a group assigned to argue against a plan or to attack a system, independently of the people who made it.

Post-mortems (Book 15’s blameless reviews) look back at an outcome; [pre-mortems](#def-fm-culture-and-decision-making-premortem) look forward at a decision. Both work only if they separate the quality of the decision from the luck of the outcome, which requires the journal.

## 26.5 How good firms disagree

A firm disagrees well when it collects views before discussing them, states them as numbers, lets a junior person’s forecast count as much as a senior’s, assigns someone to argue the other side, and scores everyone later. Book 7’s research review and pre-registration are the same discipline applied to research; the research log is a [decision journal](#def-fm-culture-and-decision-making-journal) for strategies.

**Method 26.7 (Running a betting culture).**

1. Keep a [decision journal](#def-fm-culture-and-decision-making-journal) for every significant decision, with probabilities and expected values written before the outcome.
2. Collect forecasts independently before meetings; discuss; collect again; aggregate the log-odds.
3. Score forecasters quarterly with the [Brier score](#def-fm-culture-and-decision-making-brier) and its decomposition; train calibration where the reliability term is large.
4. Review decisions by the process and the information available, not the outcome; hold [pre-mortems](#def-fm-culture-and-decision-making-premortem) for large decisions and give a [red team](#def-fm-culture-and-decision-making-premortem) the right to be heard.
5. Reward good calibration and good process, not lucky outcomes (chapter 10).

## 26.6 Build: the decision journal

**Purpose.** A [decision journal](#def-fm-culture-and-decision-making-journal) and the scoring of its forecasts: [Brier scores](#def-fm-culture-and-decision-making-brier), Murphy’s decomposition, calibration, aggregation and an outcome-bias check.

**Interface.** `firm.decisionlog`: `Entry`, `brier`, `murphy`, `calibration`, `aggregate`, `logit`, `expit`, `outcome_bias`.

**Rules.** Probabilities are recorded before outcomes; the decomposition is exact when forecasts are constant within bins; aggregation is over the same events.

**Acceptance tests.** `code/firm/decisionlog/tests/`: the decomposition’s identity on binned forecasts; the three aggregations on hand numbers; the outcome-bias count.

**Stretch.** Scoring continuous forecasts; weighting forecasters by past skill; a journal of limit and hiring decisions linked to chapter 12’s register.

Sources and further reading

- B. A. Mellers et al., “Psychological strategies for winning a geopolitical forecasting tournament”, *Psychological Science* 25(5), 2014.
- G. W. Brier, *Monthly Weather Review* 78(1), 1950; A. H. Murphy, *Journal of Applied Meteorology* 12, 1973; V. A. Satopaa et al., *International Journal of Forecasting* , 2014.
- J. Baron and J. C. Hershey, “Outcome bias in decision evaluation”, *Journal of Personality and Social Psychology* 54(4), 1988.

## 26.7 Exercises

**Exercise 26.1 ★.**

Compute the [Brier score](#def-fm-culture-and-decision-making-brier) of the forecast 0.7 for three events of which two happened. What does a constant 0.5 score?

**Solution of Exercise 26.1.**

$(0.09+0.09+0.49)/3=0.223$; a constant 0.5 scores 0.25.

**Exercise 26.2 ★.**

Average the probabilities 0.6 and 0.8, and their log-odds. Which is further from one half?

**Solution of Exercise 26.2.**

0.70 for the probabilities; the log-odds average gives 0.710, further from one half.

**Exercise 26.3 ★.**

With five forecasters whose errors are independent, by what factor does averaging cut the error variance? With a correlation of 0.5?

**Solution of Exercise 26.3.**

By 5 when independent; with $\rho=0.5$ the variance is $0.5+0.5/5=0.6$ of one forecaster’s, a factor of only 1.67.

**Exercise 26.4 ★★.**

Why does forecaster 5 score worse than forecaster 1 although its resolution is about the same?

**Solution of Exercise 26.4.**

Its reliability term is 0.040 against 0.017: it stretches its log-odds, so its extreme forecasts come true far less often than stated; its resolution is about the same.

**Exercise 26.5 ★★.**

Why does extremising hurt in the chapter’s year, and when would it help?

**Solution of Exercise 26.5.**

The forecasters are already overconfident, so pushing the average further from one half adds miscalibration. Extremising helps when forecasters hold different information and are individually well calibrated, so that their average is too moderate.

**Exercise 26.6 ★★.**

How should a desk run a meeting to decide whether a strategy is broken, given [Proposition 26.4](#prop-fm-culture-and-decision-making-rho)?

**Solution of Exercise 26.6.**

Collect each person’s probability that the strategy is broken before the meeting, discuss the evidence, collect again, and record the aggregate of the log-odds in the journal with the kill criterion of chapter 8.

**Exercise 26.7 ★★★.**

*Coding.* Score the journal’s bets by outcome and by process. What share of positive-expected-value decisions would an outcome-based review call mistakes?

**Solution of Exercise 26.7.**

Of 362 bets with positive expected value by the forecaster’s own probabilities, 123 lost: an outcome-based review would call 34% of them mistakes.

**Exercise 26.8 ★★★.**

*Find the flaw.* “We made money on that trade, so it was a good decision.”

**Solution of Exercise 26.8.**

A winning outcome says little about the decision: judge it by the probabilities and information at the time, recorded in the journal, and by the calibration of the person’s past forecasts.

## 26.8 Problem: Five Opinions or One

**Problem 26.1.**

Weekend problem — five opinions or one

A head of research wants to replace the desk’s weekly consensus meeting with a scored forecasting process.

**Part I — Principles.**

1. What did the forecasting tournament of the hook find?
2. Define a [decision journal](#def-fm-culture-and-decision-making-journal) .
3. Define the [Brier score](#def-fm-culture-and-decision-making-brier) and give its value for a constant one half.
4. State Murphy’s decomposition and what each term measures.

**Part II — The year.**

5. Give the five forecasters’ scores and the truth’s.
6. Decompose the scores of forecasters 1 and 5 and of the average.
7. Compare the three aggregation rules.
8. Read the calibration curves.

**Part III — The discussion.**

9. Define [forecast aggregation](#def-fm-culture-and-decision-making-agg) .
10. State and prove [Proposition 26.4](#prop-fm-culture-and-decision-making-rho) .
11. Give the [Brier scores](#def-fm-culture-and-decision-making-brier) of independent forecasts and of a consensus at $\rho=0.5$ and $\rho=1$ .
12. Why does discussion correlate errors?
13. How would you run the meeting instead?

**Part IV — Culture.**

14. Define [outcome bias](#def-fm-culture-and-decision-making-outcome) and give the journal’s evidence of it.
15. Define a [pre-mortem](#def-fm-culture-and-decision-making-premortem) and a [red team](#def-fm-culture-and-decision-making-premortem) .
16. How should pay reflect forecasting skill?
17. What would you record in a journal for a limit increase?
18. What could go wrong with scoring people?
19. State the *named result* : the Brier-score improvement from aggregating five independent forecasts against one discussed forecast whose errors are correlated at a given level.
20. In two sentences, write the proposal.

**Solution of Problem 26.1.**

1. Probability training, teaming and tracking improved calibration and resolution; with aggregation, the best forecasts two years in a row.
2. See [Definition 26.1](#def-fm-culture-and-decision-making-journal) .
3. See [Definition 26.2](#def-fm-culture-and-decision-making-brier) ; 0.25.
4. Reliability minus resolution plus uncertainty: miscalibration, discrimination between events, and the events’ own unpredictability.
5. 0.240, 0.225, 0.252, 0.245 and 0.262; the truth 0.214.
6. Forecaster 1: reliability 0.017, resolution 0.027; forecaster 5: 0.040 and 0.026; the log-odds average: 0.010 and 0.039, with uncertainty 0.25.
7. Probability mean 0.220, log-odds mean 0.223, extremised 0.237.
8. The overconfident forecaster’s curve is flatter than the diagonal; the average lies close to it.
9. See [Definition 26.3](#def-fm-culture-and-decision-making-agg) .
10. See [Proposition 26.4](#prop-fm-culture-and-decision-making-rho) .
11. 0.207 independent; 0.224 at 0.5; 0.237 at 1.
12. People anchor on the first number, the loudest argument or the senior person’s view.
13. Independent numbers first, discussion second, a second round, aggregation of the log-odds.
14. See [Definition 26.5](#def-fm-culture-and-decision-making-outcome) ; 34% of positive-expected-value bets lost, and forecaster 2 won 66% while forecasting 77.7%.
15. See [Definition 26.6](#def-fm-culture-and-decision-making-premortem) .
16. Through scores over long periods and many forecasts, not single outcomes (chapter 10).
17. The breach, the options, the probability that the position recovers, the expected cost of each option, who approved, and the review date.
18. Gaming by forecasting only easy events, hedging towards one half, and discouraging dissent; score all significant decisions, with calibration and resolution.
19. In the chapter’s model, aggregating five independent forecasts scores 0.207 against 0.224 for one discussed forecast at $\rho=0.5$ (0.017 better) and 0.237 at $\rho=1$ (0.030 better).
20. Collect written probabilities from each person before any discussion, aggregate their log-odds and score everyone quarterly; keep the meeting, but after the numbers are in.

## 26.9 Interview questions

**Interview question 26.1 ★ researcher.**

What is a [Brier score](#def-fm-culture-and-decision-making-brier), and what does a score of 0.25 mean?

**Solution of Interview question 26.1.**

The mean squared error of probability forecasts of binary outcomes; 0.25 is what always saying one half scores.

*What the interviewer is looking for: the no-skill benchmark.*

**Interview question 26.2 ★ trader.**

You lost money on a trade you still think was right. How do you tell?

**Solution of Interview question 26.2.**

Look at the probabilities and information recorded when you decided and at your calibration over many such trades, not at this outcome.

*What the interviewer is looking for: process over outcome.*

**Interview question 26.3 ★★ researcher.**

Why average log-odds rather than probabilities?

**Solution of Interview question 26.3.**

Log-odds add evidence symmetrically and are not dragged towards one half by a single moderate forecaster; averaging probabilities is too conservative when forecasters hold different information.

*What the interviewer is looking for: evidence adds in log-odds.*

**Interview question 26.4 ★★ researcher, trader.**

Five people’s errors are correlated at 0.5. How much does averaging their forecasts reduce the error variance?

**Solution of Interview question 26.4.**

To $0.5+0.5/5=0.6$ of one person’s: a 40% reduction, against 80% if independent.

*What the interviewer is looking for: $\rho+(1-\rho)/K$.*

**Interview question 26.5 ★★ risk.**

How would you run a [pre-mortem](#def-fm-culture-and-decision-making-premortem) for a new market entry?

**Solution of Interview question 26.5.**

Before committing, have the team assume the entry failed after a year and write down why; rank the reasons, add the likely ones to the plan’s kill criteria and critical path.

*What the interviewer is looking for: imagined failure, then action.*

**Interview question 26.6 ★★★ researcher.**

Derive Murphy’s decomposition of the [Brier score](#def-fm-culture-and-decision-making-brier).

**Solution of Interview question 26.6.**

Group forecasts into bins; expand $\frac1N\sum(p_i-y_i)^2$ around each bin’s observed frequency and around the overall frequency; with forecasts constant within bins the cross terms vanish, leaving reliability minus resolution plus uncertainty.

*What the interviewer is looking for: the binning and the vanishing cross terms.*
