---
title: "Engineering Organisation Design"
book: "The Desk and the Firm"
subject: quant
language: en
chapter: 21
exercises: 8
source: https://one-course.com/books/quant/16/en/chapter/21-engineering-organisation-design
---

# Chapter 21 — Engineering Organisation Design

In April 1968 Melvin Conway published the observation that “organizations which design systems … are constrained to produce designs which are copies of the communication structures of these organizations”. A trading firm that splits its engineers by asset class ends up with three order gateways, three risk checks and three market-data handlers, whether or not it wanted them, because three teams that rarely talk build three of everything. The organisation chart is the first draft of the architecture.

## 21.1 Embedded against central teams

**Definition 21.1 (Embedded team, central team).**

An *embedded team* sits with one business (a desk or strategy group), reports into it and builds everything that business needs. A *central team* serves several businesses with one function, such as order gateways, market data or risk checks, and reports into technology.

The two designs trade the same things. [Embedded teams](#def-fm-engineering-organisation-design-teams) talk to their traders every day, change quickly and own their whole stack; they duplicate what other desks also build, and each copy needs its own maintenance and on-call. [Central teams](#def-fm-engineering-organisation-design-teams) build one thing well for everyone and can consolidate duplicates; every change a desk needs from them crosses a team boundary. Chapter 19’s platform and [product teams](https://one-course.com/books/quant/16/en/chapter/19-technology-strategy#def-fm-technology-strategy-teams) are the modern names for the two poles, and Skelton and Pais’s *Team Topologies* (2019) their vocabulary; MacCormack, Rusnak and Baldwin (2012) tested Conway’s “mirroring” of organisation in architecture empirically.

**Definition 21.2 (Coordination cost).**

The *coordination cost* of an organisation design is the work of keeping teams in step across the dependencies between the services they own; the chapter measures it as the number of changes a week, to services that other teams’ services depend on, which those other teams must learn about, test against and schedule around.

## 21.2 The quant developer and other hybrid roles

**Definition 21.3 (Quant developer).**

A *quant developer* is an engineer who implements and runs trading models: who turns a researcher’s prototype into production code, owns its performance and correctness, and knows enough of the model to change it safely.

Hybrid roles exist because the handoff between research and production is where models break. Book 12 described the machine-learning engineer and the handoff specification between a model owner and the platform; the [quant developer](#def-fm-engineering-organisation-design-qd) is the same role for every kind of model. In an embedded design they sit on the desk; in a central design they sit between the desk and the platform, and are the people the [coordination cost](#def-fm-engineering-organisation-design-coord) falls on.

## 21.3 Ownership: every service has an owner

**Definition 21.4 (Service owner, bus factor).**

A *service owner* is the named team, with a named lead, accountable for a service’s correctness, changes, on-call and retirement. The *bus factor* of a service is the number of people who can operate and change it; the bus factor of an organisation is the smallest over its services: the number of departures that would leave some service with no one who understands it.

Ownership turns the design into data ([Listing 21.1](#lst-fm-engineering-organisation-design-measures)): a table of services with their owner, their dependencies, how often each changes and how often it pages. From it three measures follow: the [coordination cost](#def-fm-engineering-organisation-design-coord), the on-call load per engineer and the [bus factor](#def-fm-engineering-organisation-design-owner).

```python
def cross_edges(deps, design):
    return [(a, b) for a, b in deps if design.owner[a] != design.owner[b]]


def coordination(services, deps, design):
    freq = {s.name: s.change_per_week for s in services}
    return sum(freq[b] for _, b in cross_edges(deps, design))


def pages(services, design):
    load = {t: 0.0 for t in design.team_size}
    for s in services:
        load[design.owner[s.name]] += s.pages_per_week
    return {t: load[t] / design.team_size[t] for t in design.team_size}


def bus_factor(services, design):
    bf = {s.name: design.team_size[design.owner[s.name]] + design.backups.get(s.name, 0) for s in services}
    low = min(bf.values())
    return low, sorted(k for k, v in bf.items() if v == low)


def move(design, from_team, to_team):
    """One engineer moves between teams."""
    sizes = dict(design.team_size)
    sizes[from_team] -= 1
    sizes[to_team] += 1
    return replace(design, team_size=sizes)
```

***Listing 21.1.** The three measures of an organisation design: cross-team dependencies weighted by change frequency, pages per engineer by team, and the bus factor. code/firm/teamtopo/firm_teamtopo.py*

**Proposition 21.5 (Designs move load, services create it).**

For a fixed set of services and engineers, the mean on-call load per engineer is the same in every design: total pages divided by engineers. A design changes only how the load is distributed across teams, the [coordination cost](#def-fm-engineering-organisation-design-coord) and the [bus factor](#def-fm-engineering-organisation-design-owner). Only removing or merging services, or making them page less, lowers the mean.

**Proof.** Each service’s pages are counted once, in its owner’s team, whatever the owner; summing over teams gives the total pages, and dividing by the engineers, who are all in some team, gives the mean. ∎

## 21.4 Tutorial: three order gateways

**Goal.** Compare an embedded and a central design for twenty services and fifteen engineers, then consolidate, then move one engineer and see what breaks. **End state:** [Table 21.1](#tab-fm-engineering-organisation-design-designs) and the load chart ([Figure 21.2](#fig-fm-engineering-organisation-design-pages)).

1. **The services.** Three desks (equities, futures, options), each with a strategy, an order gateway, a risk check, a market-data handler and a pricing library; five shared services (security master, post-trade, monitoring, deployment, research platform); 33 dependencies; 34.5 changes and 6.2 pages a week in all ( `fm_teams` , illustrative).
2. **Embedded.** Each desk’s team of four owns its five services; an infrastructure team of three owns the shared ones.
3. **Central.** Desk teams of two own the strategies and pricing; teams of two own the three gateways, the three risk checks and the three market-data handlers; infrastructure as before.
4. **Consolidated.** The central design’s natural next step: one gateway, one risk check and one market-data handler, owned by a trading [platform team](https://one-course.com/books/quant/16/en/chapter/19-technology-strategy#def-fm-technology-strategy-teams) of three; desk teams of three; fourteen services.
5. **The move.** One engineer leaves a desk (embedded) or the gateway team (central) for infrastructure.

|  |  |  | cross-team | coordination | max pages per | bus |
| --- | --- | --- | --- | --- | --- | --- |
| design | teams | services | edges | (changes/week) | engineer-week | factor |
| embedded | 4 | 20 | 18 | 22.5 | 0.43 | 3 |
| central | 7 | 20 | 30 | 33.0 | 0.90 | 2 |
| central, consolidated | 5 | 14 | 19 | 29.0 | 0.67 | 3 |

***Table 21.1.** Three designs for fifteen engineers. Mean pages per engineer a week: 0.41 for both designs of the twenty services, 0.29 after consolidation. Data: `fm_teams.compare_all`.*

The embedded design has the least coordination (22.5 changes a week crossing teams, on 18 edges), the most even on-call load (0.43 pages per engineer a week at most) and a [bus factor](#def-fm-engineering-organisation-design-owner) of 3. Its cost is invisible in the table: it runs three of everything. Centralising the same twenty services is the worst of both: 30 cross-team edges, 33 changes a week to coordinate, the market-data team carrying 0.90 pages per engineer a week, and a [bus factor](#def-fm-engineering-organisation-design-owner) of 2, since every function now sits with two people. The mean load does not move ([Proposition 21.5](#prop-fm-engineering-organisation-design-load)): reorganising alone redistributes work and adds boundaries.

Consolidation is what the central design is for. Merging the three gateways, risk checks and market-data handlers leaves fourteen services and 4.3 pages a week, 31% fewer; the mean load falls to 0.29 pages per engineer, the [bus factor](#def-fm-engineering-organisation-design-owner) returns to 3 with teams of three, and the desks gain an engineer each. Coordination stays higher than in the embedded design (29.0 changes a week), because every desk now depends on one shared gateway: that is the price of having one.

![Coordination under the three designs: the number of dependencies that cross a team boundary, and the changes a week that cross them. Centralising without consolidating adds boundaries; consolidating removes some of them again. Data: fm_teams.compare_all.](https://one-course.com/images/onecourse/chapters/quant-16/fm-engineering-organisation-design/fig-0bddcdb444d1.svg)

***Figure 21.1.** Coordination under the three designs: the number of dependencies that cross a team boundary, and the changes a week that cross them. Centralising without consolidating adds boundaries; consolidating removes some of them again. Data: `fm_teams.compare_all`.*

## 21.5 On-call and its cost

On-call load is measured in pages, but it is paid in hours. Google’s site-reliability book caps on-call at a quarter of an engineer’s time and finds that an incident, with its root-cause analysis, remediation and follow-up, takes about six hours; at most two such incidents fit in a twelve-hour shift ([Box 21.1](#dat-fm-engineering-organisation-design-sre)). On the chapter’s numbers the central design’s market-data team spends $0.9\times6=5.4$ hours a week each on incidents, 13.5% of a forty-hour week; after one engineer leaves the gateway team, its remaining engineer carries 1.5 pages a week, nine hours, 22.5% of the week and close to the cap ([Figure 21.2](#fig-fm-engineering-organisation-design-pages)).

**As of September 2026 — A published on-call budget.**

Google’s *Site Reliability Engineering* (2016), chapter “Being On-Call”: at least 50% of site-reliability engineers’ time goes to engineering and no more than 25% to on-call; handling an incident takes about six hours on average, so the maximum is two incidents per twelve-hour shift. Book 15 applies these budgets to a trading platform’s on-call rotation.

![On-call load per engineer by team under the three designs (the three desk teams carry the same load in each design). The central design concentrates pages on the small function teams. Data: firm.teamtopo.pages.](https://one-course.com/images/onecourse/chapters/quant-16/fm-engineering-organisation-design/fig-1972c152d5f2.svg)

***Figure 21.2.** On-call load per engineer by team under the three designs (the three desk teams carry the same load in each design). The central design concentrates pages on the small function teams. Data: `firm.teamtopo.pages`.*

The move shows the [bus factor](#def-fm-engineering-organisation-design-owner) at work. In the embedded design, moving one engineer off the equities desk leaves its five services with three people and raises that desk’s load to 0.57 pages a week each; nothing is orphaned. In the central design, moving one engineer off the gateway team leaves the three gateways with one person each: the [bus factor](#def-fm-engineering-organisation-design-owner) falls to 1, and one resignation would leave the firm unable to change how it sends orders. In the consolidated design the same move leaves the trading platform with two people and a [bus factor](#def-fm-engineering-organisation-design-owner) of 2.

![Conway’s law in a trading firm: three desk teams build three gateways (left); a platform team builds one that all three desks depend on (dashed, right). The second design has fewer services and more cross-team dependencies. Schematic.](https://one-course.com/images/onecourse/chapters/quant-16/fm-engineering-organisation-design/fig-412e16580a6a.svg)

***Figure 21.3.** Conway’s law in a trading firm: three desk teams build three gateways (left); a [platform team](https://one-course.com/books/quant/16/en/chapter/19-technology-strategy#def-fm-technology-strategy-teams) builds one that all three desks depend on (dashed, right). The second design has fewer services and more cross-team dependencies. Schematic.*

## 21.6 Measuring an engineering organisation

The three measures are cheap to keep once ownership is data: [coordination cost](#def-fm-engineering-organisation-design-coord) from the dependency graph and change logs, on-call load from the paging system, [bus factor](#def-fm-engineering-organisation-design-owner) from who has changed and operated each service in the last year. Forsgren, Humble and Kim’s *Accelerate* (2018) proposes delivery measures (how often teams deploy, how long a change takes to reach production, how often changes fail, how fast service is restored) that complete them.

**Method 21.6 (Designing the engineering organisation).**

1. Keep the service catalogue as data: owner, dependencies, change frequency, page rate, who can operate it.
2. Draw teams around the communication the work needs (Conway’s criterion): embed what changes with the desk, centralise what several desks need in the same form, and consolidate what is duplicated.
3. Measure coordination, on-call load against a budget, and the [bus factor](#def-fm-engineering-organisation-design-owner) , before and after every reorganisation.
4. Keep every service’s [bus factor](#def-fm-engineering-organisation-design-owner) at three or more; pair new people on the weakest services first.
5. Review the design when a desk or a venue is added: each is a new service or a new dependency.

## 21.7 Build: the organisation as a graph

**Purpose.** Services, dependencies, change and incident rates and team assignments as data, with the measures of a design.

**Interface.** `firm.teamtopo`: `Service`, `Design`, `cross_edges`, `coordination`, `pages`, `bus_factor`, `move`, `report`.

**Rules.** Every service has exactly one owning team; a service’s [bus factor](#def-fm-engineering-organisation-design-owner) counts its team and named backups; the mean load is invariant under reassignment.

**Acceptance tests.** `code/firm/teamtopo/tests/`: cross-team edges and coordination on a three-service graph; pages by team; the [bus factor](#def-fm-engineering-organisation-design-owner) before and after a move.

**Stretch.** Knowledge measured from commit and deployment history; the cost of a reorganisation itself; optimising the assignment for a weighted sum of the three measures.

Sources and further reading

- M. E. Conway, “How do committees invent?”, *Datamation* , April 1968.
- A. MacCormack, C. Baldwin and J. Rusnak, *Research Policy* 41(8), 2012; M. Skelton and M. Pais, *Team Topologies* , 2019.
- B. Beyer et al. (eds), *Site Reliability Engineering* , 2016; N. Forsgren, J. Humble and G. Kim, *Accelerate* , 2018.

## 21.8 Exercises

**Exercise 21.1 ★.**

State Conway’s thesis and his criterion for organising a design effort.

**Solution of Exercise 21.1.**

Organisations that design systems are constrained to produce designs that copy their communication structures; a design effort should therefore be organised according to the need for communication.

**Exercise 21.2 ★.**

What is the mean on-call load per engineer in each design, and why is it the same for the embedded and central designs?

**Solution of Exercise 21.2.**

$6.2/15=0.41$ pages a week in both designs of the twenty services, and $4.3/15=0.29$ after consolidation; by [Proposition 21.5](#prop-fm-engineering-organisation-design-load) reassignment cannot change the mean.

**Exercise 21.3 ★.**

How many hours a week does the central market-data team spend on incidents, at six hours an incident, and what share of the on-call cap is it?

**Solution of Exercise 21.3.**

$0.9\times6=5.4$ hours a week each: 13.5% of a forty-hour week, just over half of the 25% cap.

**Exercise 21.4 ★★.**

Why does centralising the same twenty services raise coordination from 22.5 to 33.0 changes a week?

**Solution of Exercise 21.4.**

Each desk’s strategy now depends on gateways, market data and pricing inputs owned by other teams: the edges from strategies to gateways and market data, and from gateways to risk checks, all cross boundaries, and they carry frequent changes.

**Exercise 21.5 ★★.**

What happens to the [bus factor](#def-fm-engineering-organisation-design-owner) in each design when one engineer moves, and which services are at risk?

**Solution of Exercise 21.5.**

Embedded: it stays 3 (the equities desk keeps three people). Central: it falls to 1, and the three gateways each depend on one person. Consolidated: it falls to 2, for the one gateway, risk check and market-data handler.

**Exercise 21.6 ★★.**

Where should the [quant developers](#def-fm-engineering-organisation-design-qd) sit in each design, and why?

**Solution of Exercise 21.6.**

Embedded: on the desks, where the models change. Central: between the desks and the platform, as the owners of the handoff specifications, since the [coordination cost](#def-fm-engineering-organisation-design-coord) falls on them.

**Exercise 21.7 ★★★.**

*Coding.* In the consolidated design, move one engineer from the trading platform to infrastructure. What are the [bus factor](#def-fm-engineering-organisation-design-owner) and the largest load?

**Solution of Exercise 21.7.**

A [bus factor](#def-fm-engineering-organisation-design-owner) of 2 for the gateway, the risk check and the market-data handler, and a load of 1.0 page a week on each of the two remaining platform engineers.

**Exercise 21.8 ★★★.**

*Find the flaw.* “Centralising our engineers will cut our on-call load.”

**Solution of Exercise 21.8.**

Moving the same services between teams leaves the mean load unchanged and concentrates it on small function teams; only consolidating or fixing the noisy services lowers it.

## 21.9 Problem: Three Order Gateways

**Problem 21.1.**

Weekend problem — three order gateways

A new chief technology officer inherits three desks with their own engineers and three of everything, and must propose an organisation.

**Part I — The principles.**

1. State Conway’s thesis.
2. Define embedded and [central teams](#def-fm-engineering-organisation-design-teams) and their trade-off.
3. Define a [quant developer](#def-fm-engineering-organisation-design-qd) and say why the role exists.
4. Define [coordination cost](#def-fm-engineering-organisation-design-coord) , [service owner](#def-fm-engineering-organisation-design-owner) and [bus factor](#def-fm-engineering-organisation-design-owner) .

**Part II — The model.**

5. Describe the services and their dependencies.
6. State and prove [Proposition 21.5](#prop-fm-engineering-organisation-design-load) .
7. Give the three measures for the embedded and central designs.
8. Give them for the consolidated design and explain what consolidation buys and costs.

**Part III — Stress.**

9. What is the published on-call budget, and how close does each design come to it?
10. Move one engineer in each design and give the result.
11. Which services should be paired first, and why?
12. How would you measure the [bus factor](#def-fm-engineering-organisation-design-owner) from data?

**Part IV — The proposal.**

13. Which services should stay embedded, which be centralised, which be consolidated?
14. How would you sequence the change?
15. What delivery measures would you add?
16. What is the risk of a single shared gateway, and how would you manage it?
17. How often should the design be reviewed?
18. Why is reorganising without consolidating the worst option in the model?
19. State the *named result* : [coordination cost](#def-fm-engineering-organisation-design-coord) and pages per engineer per week under the embedded and the central design, and the [bus factor](#def-fm-engineering-organisation-design-owner) of each.
20. In two sentences, write the proposal.

**Solution of Problem 21.1.**

1. See [Exercise 21.1](#exo-fm-engineering-organisation-design-1) .
2. See [Definition 21.1](#def-fm-engineering-organisation-design-teams) : closeness and speed against duplication, and one good copy against boundaries.
3. See [Definition 21.3](#def-fm-engineering-organisation-design-qd) ; the handoff from research to production is where models break.
4. See Definitions [21.2](#def-fm-engineering-organisation-design-coord) and [21.4](#def-fm-engineering-organisation-design-owner) .
5. Three desks with five services each, five shared services, 33 dependencies, 34.5 changes and 6.2 pages a week.
6. See [Proposition 21.5](#prop-fm-engineering-organisation-design-load) .
7. Embedded: 22.5 changes a week, 0.43 pages per engineer at most, [bus factor](#def-fm-engineering-organisation-design-owner) 3. Central: 33.0, 0.90, 2.
8. 29.0, 0.67 and 3, with a mean load of 0.29; one copy of each function and a free engineer per desk, at the cost of every desk depending on one team.
9. At most 25% of time on-call, two incidents per twelve-hour shift; the busiest [central team](#def-fm-engineering-organisation-design-teams) uses 13.5% of a week, the embedded desks about 6%.
10. See [Exercise 21.5](#exo-fm-engineering-organisation-design-5) .
11. The services at the minimum [bus factor](#def-fm-engineering-organisation-design-owner) : the gateways in the central design.
12. From who has changed, deployed or handled incidents on each service in the last year.
13. Embed strategies and pricing; consolidate and centralise gateways, risk checks and market data; keep shared services central.
14. Consolidate one function at a time, with its team formed first and the old copies retired only after the new one has run in parallel.
15. Deployment frequency, lead time for changes, change failure rate, time to restore service.
16. One failure or one departure stops every desk; keep its [bus factor](#def-fm-engineering-organisation-design-owner) at three or more, test changes with canaries, and hold a tested fallback.
17. At each new desk, venue or major service, and yearly.
18. It adds boundaries and concentrates load without removing any service.
19. Embedded: 22.5 changes a week, 0.43 pages per engineer a week, [bus factor](#def-fm-engineering-organisation-design-owner) 3; central: 33.0, 0.90, 2.
20. Keep strategy work embedded and consolidate the duplicated gateways, risk checks and market-data handlers under one [platform team](https://one-course.com/books/quant/16/en/chapter/19-technology-strategy#def-fm-technology-strategy-teams) of at least three; measure coordination, on-call load and [bus factor](#def-fm-engineering-organisation-design-owner) before and after every step.

## 21.10 Interview questions

**Interview question 21.1 ★ developer.**

What is Conway’s law, and where have you seen it?

**Solution of Interview question 21.1.**

Systems copy the communication structure of the organisations that build them; for example one order gateway per desk where desks do not share engineers.

*What the interviewer is looking for: the thesis and a concrete instance.*

**Interview question 21.2 ★ developer, researcher.**

What does a [quant developer](#def-fm-engineering-organisation-design-qd) do that a researcher and a software engineer do not?

**Solution of Interview question 21.2.**

Turns models into production code, owns their performance and correctness, and understands the model well enough to change it safely.

*What the interviewer is looking for: the handoff.*

**Interview question 21.3 ★★ developer.**

How would you compute the [bus factor](#def-fm-engineering-organisation-design-owner) of a codebase?

**Solution of Interview question 21.3.**

For each file or service, count the people who have changed it meaningfully in a recent window; the [bus factor](#def-fm-engineering-organisation-design-owner) is the smallest set of people whose removal leaves some part with nobody, approximated by the minimum count.

*What the interviewer is looking for: history-based knowledge.*

**Interview question 21.4 ★★ developer.**

Your team gets two pages a night. What do you do?

**Solution of Interview question 21.4.**

Measure which services page and why, fix or silence the noisiest, add runbooks and automation, and rebalance the rotation; two a night is above a sustainable budget.

*What the interviewer is looking for: fix the source, then the rota.*

**Interview question 21.5 ★★ trader, developer.**

Should each trading desk have its own engineers?

**Solution of Interview question 21.5.**

For what makes the desk different, yes; for what every desk needs in the same form, no: shared platforms avoid three copies.

*What the interviewer is looking for: the embedded-central split by service.*

**Interview question 21.6 ★★★ developer, researcher.**

Formulate the assignment of services to teams as an optimisation problem.

**Solution of Interview question 21.6.**

Choose an owner for each service and a team for each engineer to minimise a weighted sum of coordination (cross-team edges weighted by change frequency), on-call load above a budget and a penalty for [bus factors](#def-fm-engineering-organisation-design-owner) below a floor; an integer programme.

*What the interviewer is looking for: the three measures as objective and constraints.*
