---
title: "Point-in-Time Data and the Biases"
book: "Research Craft: Predictors, Backtests, Measurement, Portfolios"
subject: quant
language: en
chapter: 3
exercises: 8
source: https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases
---

# Chapter 3 — Point-in-Time Data and the Biases

On 3 October 2008 the US employment report said that the economy had lost 159 000 jobs in September. Today’s database says 451 000. Between the two, the figure was revised six times, three of them by the annual benchmark revisions that the statistical agency publishes each February. A macro signal backtested on today’s database would have traded on a number nobody had in October 2008, and would have seen a recession deeper and clearer than the one anyone saw at the time. This chapter is about the difference between what is true now and what was known then: [look-ahead bias](#def-rs-point-in-time-data-and-the-biases-lookahead) and the time stamps that prevent it, [survivorship bias](#def-rs-point-in-time-data-and-the-biases-survivorship), revisions and [restatements](#def-rs-point-in-time-data-and-the-biases-vintage), the storage that keeps every vintage of a fact, and the [as-of join](#def-rs-point-in-time-data-and-the-biases-asof) that reads it. Its data are the real-time vintages of US payrolls and a simulated stock market with delistings.

## 3.1 Look-ahead bias and timestamp semantics

**Definition 3.1 (Look-ahead bias, point-in-time data).**

A backtest has *look-ahead bias* when a decision at time $t$ uses information that was not available at $t$. *Point-in-time data* are data stored so that, for any past instant, they return exactly what was known at that instant: the values then published, the universe then listed, the identifiers then used.

Look-ahead rarely comes from a deliberate peek. It comes from a time stamp that means something other than what the code assumes. A daily [bar](https://one-course.com/books/quant/7/en/chapter/2-market-data-for-research#def-rs-market-data-for-research-bar) stamped with its date is known only at the close, so a signal computed from it cannot trade at that day’s close. An earnings figure stamped with its fiscal quarter’s end is known weeks later, on the filing date (chapter 11). An index change stamped with its effective date was announced days before (One Quant Book 1, chapter 15), and the announcement is when the trade starts. A trade stamped by the exchange arrives at the firm later, and the arrival time is the one the firm could act on (Book 1, chapter 28). Each fact therefore needs two times.

**Definition 3.2 (Valid time, knowledge time, bitemporal data).**

The *valid time* of a value is the time it describes: September 2008 for September’s payroll change, the fourth quarter for a quarterly earnings figure. Its *knowledge time* is when it became available to the user: the release, the filing, the vendor’s delivery, the arrival of the message. Data are *bitemporal* when every value is stored with both, and a new value for the same valid time is added rather than overwriting the old one.

![Bitemporal data. Each dot is a value stored for a month (valid time) as published at a date (knowledge time); a month first appears the month after it and is revised later. A backtest must read horizontal cuts at its decision times, never the top one; the vertical line through September is the life of one number, drawn in .](https://one-course.com/images/onecourse/chapters/quant-7/rs-point-in-time-data-and-the-biases/fig-96b21080a761.svg)

***Figure 3.1.** [Bitemporal data](#def-rs-point-in-time-data-and-the-biases-bitemporal). Each dot is a value stored for a month ([valid time](#def-rs-point-in-time-data-and-the-biases-bitemporal)) as published at a date ([knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal)); a month first appears the month after it and is revised later. A backtest must read horizontal cuts at its decision times, never the top one; the vertical line through September is the life of one number, drawn in [Figure 3.3](#fig-rs-point-in-time-data-and-the-biases-life).*

## 3.2 Survivorship bias

**Definition 3.3 (Survivorship bias).**

*Survivorship bias* is the error of studying only the securities, funds or firms that exist at the end of the sample: those that were delisted, liquidated or merged away are missing, and they are not missing at random.

A stock universe taken from today’s list of listed companies has already dropped every company that failed. A database of funds built from the funds reporting today has dropped those that closed. The dropped names are, on average, the ones that did badly, so their absence flatters every average. Brown, Goetzmann, Ibbotson and Ross (1992) showed how survivorship alone can create apparent persistence in fund performance; for stocks, Shumway (1997) found that the delisting returns of stocks delisted for poor performance were mostly missing from the standard US database, and that, where he could recover them from over-the-counter prices, they averaged $-30\%$.

**Proposition 3.4 (How much the survivors overstate).**

Pool all the name-months of a universe. Let a fraction $w$ of them belong to names that are delisted before the end of the sample, with average return $r_F$, and the rest to names still listed at the end, with average return $r_S$. The pooled average return of the whole universe is $(1 - w)r_S + wr_F$, so an average over the survivors overstates it by $w(r_S - r_F)$.

**Proof.** The pooled average is the weighted average of the two groups’ averages, with weights equal to their shares of name-months. ∎

The identity says where the bias comes from: not only the delisting return, a single month, but every month a failing name spends declining towards its delisting. The chapter’s simulated market makes it concrete. Two thousand stocks follow random walks in log price, with an 8% drift and 40% volatility a year; a stock whose cumulative log return falls below $-1.6$ (an 80% loss) is delisted with a return of $-30\%$ that month and replaced by a new listing. Over twenty years 2.2% of names are delisted each year. The equal-weighted universe, delistings included, returns 7.7% a year; the universe of today’s survivors, backfilled over their own histories, returns 12.9%: a bias of 5.2 points a year ([Figure 3.2](#fig-rs-point-in-time-data-and-the-biases-surv)). In the proposition’s terms, 20.9% of the name-months belong to names that later fail, with $r_F = -11.1\%$ a year against $r_S = 12.6\%$, so $w(r_S - r_F) = 5.0$ points; the remaining difference comes from averaging month by month over a changing number of survivors rather than pooling. Setting the delisting return to zero lowers the bias only to 4.5 points: the decline before delisting is most of it.

![Equal-weighted growth of one unit in a simulated market of 2 000 stocks over twenty years, with 2.2% of names delisted each year at -30\%: the investable universe ends at 4.6, the universe of the names still listed at the end at 12.9. Data: the chapter’s module, seeded.](https://one-course.com/images/onecourse/chapters/quant-7/rs-point-in-time-data-and-the-biases/fig-4f49f4c33037.svg)

***Figure 3.2.** Equal-weighted growth of one unit in a simulated market of 2 000 stocks over twenty years, with 2.2% of names delisted each year at $-30\%$: the investable universe ends at 4.6, the universe of the names still listed at the end at 12.9. Data: the chapter’s module, seeded.*

## 3.3 Restatements, revisions and vintages

**Definition 3.5 (Restatement, data vintage, backfill bias).**

A *restatement* is a later published change to a value already released: a company restating an earlier quarter’s accounts, a statistical agency revising a figure. A *data vintage* is the whole history of a series as it was published at one [knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal). *Backfill bias* arises when history is added to a database after the fact, typically when a fund or a vendor’s new coverage enters with its past returns: the entrants chose to enter because their past was good, and the backfilled history was never available in real time.

The US payroll series shows how large revisions are for a headline number. The Philadelphia Fed’s Real-Time Data Set for Macroeconomists (Croushore and Stark, 2001) keeps every vintage of the series since 1964; vintage $m$ is the history as published in the report of month $m$. For the 318 months from January 2000 to June 2026 with a first release, the monthly change was later revised by 74 000 jobs on average in absolute value, and by more than 100 000 in 24% of months; the revisions average almost zero ($-3\,000$) but have a standard deviation of 106 000. Eighteen months changed sign. The largest revision is March 2020 ($-701\,000$ first, $-1\,398\,000$ today).

The September 2008 figure shows a single number’s life ([Figure 3.3](#fig-rs-point-in-time-data-and-the-biases-life)): $-159\,000$, then $-284\,000$ and $-403\,000$ in the next two reports, $-321\,000$ after the benchmark revision of February 2009, $-458\,000$ after that of February 2010, $-434\,000$ after that of February 2011, and $-451\,000$ today. The recession as a whole was invisible in its true size while it happened ([Figure 3.4](#fig-rs-point-in-time-data-and-the-biases-y2008)): the first releases of the twelve months of 2008 sum to a loss of 1.88 million jobs, the January 2009 report put the year at 2.59 million, and today’s data say 3.55 million. A model fitted on today’s data learns a relation between economic news and markets that no participant could have acted on.

![The US payroll change for September 2008 in each of the 36 monthly vintages after it: -159, -284, -403, then the benchmark revisions of February 2009 (-321), February 2010 (-458) and February 2011 (-434); -451 in today’s vintage (dashed). Data: Federal Reserve Bank of Philadelphia, Real-Time Data Set for Macroeconomists (BLS data).](https://one-course.com/images/onecourse/chapters/quant-7/rs-point-in-time-data-and-the-biases/fig-fb332aff574b.svg)

***Figure 3.3.** The US payroll change for September 2008 in each of the 36 monthly vintages after it: $-159$, $-284$, $-403$, then the benchmark revisions of February 2009 ($-321$), February 2010 ($-458$) and February 2011 ($-434$); $-451$ in today’s vintage (dashed). Data: Federal Reserve Bank of Philadelphia, Real-Time Data Set for Macroeconomists (BLS data).*

![Monthly US payroll changes in 2008–2009 as first published and as known today. The first releases understated the losses of 2008 by 1.7 million jobs in total. Data: Federal Reserve Bank of Philadelphia, Real-Time Data Set for Macroeconomists (BLS data).](https://one-course.com/images/onecourse/chapters/quant-7/rs-point-in-time-data-and-the-biases/fig-2691495ac824.svg)

***Figure 3.4.** Monthly US payroll changes in 2008–2009 as first published and as known today. The first releases understated the losses of 2008 by 1.7 million jobs in total. Data: Federal Reserve Bank of Philadelphia, Real-Time Data Set for Macroeconomists (BLS data).*

Company data behave the same way. Accounts are restated; data vendors correct errors, change definitions and extend their coverage backwards. A vendor that adds a thousand small companies with ten years of history has created a backfilled universe, and a strategy backtested on it trades companies the firm could not have known it would cover. The cure is the same in every case: keep the vintages.

## 3.4 As-of joins and bitemporal storage

**Definition 3.6 (As-of join).**

An *as-of join* matches each row of a left table (decision times) with the last row of a right table (observations) whose time is at or before the left row’s time, optionally within a tolerance; strictly before when the right table’s rows become known only after their stamp.

An ordinary join on equal time stamps fails in both directions: it drops the decision times at which no observation has exactly the same stamp, and it happily matches a decision with an observation stamped at the same instant that arrived later. The [as-of join](#def-rs-point-in-time-data-and-the-biases-asof), on [knowledge times](#def-rs-point-in-time-data-and-the-biases-bitemporal), answers the only question a backtest may ask: what was the latest thing known?

**Method 3.7 (Making a backtest point-in-time).**

1. Give every input a [knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal) : the release or filing time, the arrival time of a message, the end of the [bar](https://one-course.com/books/quant/7/en/chapter/2-market-data-for-research#def-rs-market-data-for-research-bar) plus the processing delay. Where only a [valid time](#def-rs-point-in-time-data-and-the-biases-bitemporal) exists, derive a conservative [knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal) from a documented lag.
2. Store every vintage ( [bitemporal data](#def-rs-point-in-time-data-and-the-biases-bitemporal) ); never overwrite a value.
3. Read inputs through a guard that knows the decision time and refuses any later [knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal) .
4. Build universes, identifiers and adjustments as of the decision time as well (chapter 4).
5. Test for leakage: perturb every value known after the decision time and check that no decision changes.

## 3.5 A catalogue of the other biases

Look-ahead, survivorship and backfill are the biases of data; others come from the researcher. *Selection of the sample period*: starting a test after a crash, or ending it before one, is a trial the log must count (chapter 1). *Mismatched calendars*: a US close and a Tokyo close stamped with the same date are fifteen hours apart. *Stale prices*: an illiquid instrument’s last trade may be days old, so a portfolio marked with it looks smoother than it is (chapter 22). *Corporate actions applied with hindsight*: a price series adjusted for a split with today’s factor carries information about the split into the past (chapter 4). *Index membership*: testing a strategy on today’s index members is [survivorship bias](#def-rs-point-in-time-data-and-the-biases-survivorship) under another name. All are look-ahead in a broad sense, and all are prevented by the same discipline: at every decision time, only what was known then.

## 3.6 Tutorial: payrolls as they were known

**Goal.** Load the 2008–2009 payroll vintages into a [bitemporal](#def-rs-point-in-time-data-and-the-biases-bitemporal) store, ask what was known when, and let a guard refuse the future. **End state:** Figures [3.3](#fig-rs-point-in-time-data-and-the-biases-life) and [3.4](#fig-rs-point-in-time-data-and-the-biases-y2008); the 2008 total as known in January 2009 ($-2.59$ million) and today ($-3.55$ million).

1. **The store.** Each (entity, field, [valid time](#def-rs-point-in-time-data-and-the-biases-bitemporal)) keeps a sorted list of ([knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal), value); a query finds the last [knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal) at or before the one asked. `def put (self , entity, field, valid, known, value) -> None : rows = self ._v.setdefault((entity, field, valid), []) keys = [k for k, _ in rows] i = bisect.bisect_right(keys, known) if i > 0 and keys[i - 1 ] == known: rows[i - 1 ] = (known, value) # a correction at the same knowledge time replaces else : rows.insert(i, (known, value)) def vintages (self , entity, field, valid) -> list : return list (self ._v.get((entity, field, valid), [])) def asof (self , entity, field, valid, known): rows = self ._v.get((entity, field, valid)) if not rows: return None i = bisect.bisect_right([k for k, _ in rows], known) return rows[i - 1 ][1 ] if i else None` **Listing 3.1.** Recording and reading a value by valid and knowledge time. code/firm/pit/firm_pit.py
2. **The guard.** A view of the store for one decision time; any query about later knowledge raises. `class Guard : """A view of a store for one decision time: every query that names a later knowledge time raises.""" def __init__(self , store: Store, decision_time): self .store, self .t = store, decision_time def _check (self , known): if known > self .t: raise LookAheadError(f " asked for knowledge at { known!r} after the decision time { self .t!r} " ) def asof (self , entity, field, valid, known=None ): known = self .t if known is None else known self ._check(known) return self .store.asof(entity, field, valid, known)` **Listing 3.2.** The look-ahead guard. code/firm/pit/firm_pit.py
3. **Ask.** `real_time_view` for September 2008 with the decision time October 2008 returns $-159$ , and with September 2011 $-434$ ; a guard with the decision time November 2008 raises `LookAheadError` when asked for knowledge of December 2008. Sum the twelve months of 2008 as known in each vintage and draw the result against the vintage date.

**What to change next.** Rebuild the data from the source with `rs_fetch_payrolls.py`; add the unemployment rate, whose vintages the same data set provides, and find the months whose sign of change flipped.

## 3.7 Build: the point-in-time store

**Purpose.** Every fact the miniature firm uses in research is read through this store, so that a backtest sees only what was known at each decision time. Chapter 4 builds the security master on it, chapter 11 the fundamental features.

**Interface.** `Store.put(entity, field, valid, known, value)`; `asof(entity, field, valid, known)`; `first`, `vintages`, `snapshot(entity, field, known)`, `latest`; `Guard(store, decision_time)` with the same queries; `asof_join(left_t, right_t, right_v, tolerance, strict)`.

**Rules.** Append-only by [knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal) (a correction at an existing [knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal) replaces it); times of any comparable type used consistently; the guard raises rather than returning `None`, so a leak is loud.

**Acceptance tests.** `code/firm/pit/tests/`: the life of the September 2008 figure; a snapshot equals a vintage and `latest` equals today’s view; the guard refuses a later [knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal); the [as-of join](#def-rs-point-in-time-data-and-the-biases-asof) with and without strictness and tolerance.

**Stretch.** A columnar version keyed by (entity, [valid time](#def-rs-point-in-time-data-and-the-biases-bitemporal)) for millions of rows; storing the [knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal) of a correction separately from that of the original value; an audit report of every query a backtest made.

Sources and further reading

- D. Croushore and T. Stark, “A real-time data set for macroeconomists”, *Journal of Econometrics* 105(1), 2001; data: Federal Reserve Bank of Philadelphia, Real-Time Data Set for Macroeconomists, nonfarm payroll employment.
- T. Shumway, “The delisting bias in CRSP data”, *Journal of Finance* 52(1), 1997.
- S. J. Brown, W. Goetzmann, R. G. Ibbotson and S. A. Ross, “Survivorship bias in performance studies”, *Review of Financial Studies* 5(4), 1992.
- E. J. Elton, M. J. Gruber and C. R. Blake, “Survivor bias and mutual fund performance”, *Review of Financial Studies* 9(4), 1996.
- W. Fung and D. A. Hsieh, “Performance characteristics of hedge funds and commodity funds: natural vs. spurious biases”, *Journal of Financial and Quantitative Analysis* 35(3), 2000.
- R. W. Banz and W. J. Breen, “Sample-dependent results using accounting and market data: some evidence”, *Journal of Finance* 41(4), 1986.

## 3.8 Exercises

**Exercise 3.1 ★.**

A daily [bar](https://one-course.com/books/quant/7/en/chapter/2-market-data-for-research#def-rs-market-data-for-research-bar) is stamped with its date and holds the close. A signal uses the close of day $t$ and the backtest trades at the close of day $t$. What is wrong, and what is the earliest trade the signal allows?

**Solution of Exercise 3.1.**

The close of day $t$ is known only when the day ends, so trading at that same close uses it before it exists (unless the signal is computed from prices a few minutes earlier and sent into the closing auction). The earliest honest trade is the next session’s open, or the next close in a close-to-close test.

**Exercise 3.2 ★.**

In a universe, 15% of name-months belong to names that are later delisted, with an average return of $-20\%$ a year, and the others average 10%. By how much does a survivors-only average overstate the universe’s?

**Solution of Exercise 3.2.**

$w(r_S - r_F) = 0.15 \times (0.10 + 0.20) = 0.045$: 4.5 points a year.

**Exercise 3.3 ★.**

Using the vintages in [Figure 3.3](#fig-rs-point-in-time-data-and-the-biases-life), what did a desk know about September 2008 in January 2009, in June 2009 and in September 2011?

**Solution of Exercise 3.3.**

In January 2009 (four months after), $-403\,000$; in June 2009, after the February benchmark, $-321\,000$; in September 2011, $-434\,000$.

**Exercise 3.4 ★★.**

Decision times are 10:00:00.0, 10:00:01.0 and 10:00:02.0; quotes are stamped 09:59:59.5, 10:00:01.0 and 10:00:01.8, and each arrives 0.3 seconds after its stamp. Which quote does an [as-of join](#def-rs-point-in-time-data-and-the-biases-asof) on stamps give each decision, and which one on arrival times?

**Solution of Exercise 3.4.**

On stamps: the quotes stamped 09:59:59.5, 10:00:01.0 and 10:00:01.8. On arrival times (09:59:59.8, 10:00:01.3, 10:00:02.1): the first quote for the first two decisions and the second for the third. The join on stamps lets the second decision use a quote 0.3 seconds before it existed at the firm.

**Exercise 3.5 ★★.**

Payroll revisions from first release to today have a standard deviation of 106 000 and a mean of $-3\,000$. A surprise model reads the first release against a forecast whose error, in final data, has a standard deviation of 90 000. What share of the apparent variance of first-release surprises is revision noise, if the two are independent?

**Solution of Exercise 3.5.**

The first-release surprise is the final-data surprise minus the revision, so its variance is $90^2 + 106^2 = 19\,336$ (thousand jobs squared), of which $11\,236$, or 58%, is revision noise. A model fitted on final data sees a much cleaner signal than the one the market traded.

**Exercise 3.6 ★★.**

A hedge-fund database adds funds with their past three years of returns when they join. Funds join when their past three years beat the median. Explain the bias in the database’s average return and how to remove it.

**Solution of Exercise 3.6.**

The history before a fund joins was never available in real time and is selected: only funds with good past returns join, so their backfilled years raise the database average ([backfill bias](#def-rs-point-in-time-data-and-the-biases-vintage)). Remove it by using each fund’s returns only from its date of entry into the database (its [knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal)), or by dropping a fixed initial period of every fund.

**Exercise 3.7 ★★★.**

*Coding.* With `survivorship()`, measure the bias for barriers of $-1.2$ and $-2.0$ and for delisting returns of $0$ and $-100\%$. Which parameter moves the bias more, and why?

**Solution of Exercise 3.7.**

Barrier $-1.2$: 3.5% of names delisted a year, bias 7.0 points; barrier $-2.0$: 1.4% and 3.8 points. Delisting return 0: 4.5 points; $-100\%$: 6.7 points. The barrier moves the bias more, because it sets how many names fail and how long they spend declining; the delisting return affects only one month of each failure.

**Exercise 3.8 ★★★.**

*Find the flaw.* “We backtested our value strategy on the current members of the S&P 500 since 2000, using each company’s fundamentals as they appear in our vendor’s database today. It beats the index by 3% a year.”

**Solution of Exercise 3.8.**

Three biases at once. Survivorship: today’s members exclude the companies that failed or shrank out of the index since 2000. Index-membership look-ahead: many of today’s members were added because they had risen, a fact unknown in 2000. [Restatements](#def-rs-point-in-time-data-and-the-biases-vintage) and vendor corrections: today’s fundamentals are not those published at the time. The test needs the index membership as of each date, delisted companies with their delisting returns, and point-in-time fundamentals.

## 3.9 Problem: The Survivors’ Premium

**Problem 3.1.**

Weekend problem — how much does a list of today’s stocks flatter the past?

The chapter’s simulated market: 2 000 names, random walks in log price with an 8% drift and 40% volatility a year, delisting at a cumulative log return of $-1.6$ with a return of $-30\%$, replacement by a new listing, twenty years, seed 11.

**Part I — The market.**

1. What price loss does a cumulative log return of $-1.6$ represent?
2. What fraction of names is delisted each year?
3. What is the annual equal-weighted return of the investable universe, delistings included?
4. And of today’s survivors, over their own histories?
5. What is the bias, in points a year?

**Part II — Where it comes from.**

6. What share of name-months belongs to names that are later delisted?
7. What are $r_F$ and $r_S$ ?
8. What does [Proposition 3.4](#prop-rs-point-in-time-data-and-the-biases-survivors) give, and why does it differ from Question 5?
9. What is the bias with a delisting return of zero?
10. What is it with a delisting return of $-100\%$ ?

**Part III — Sensitivity.**

11. What are the delisting rate and the bias with a barrier of $-1.2$ ?
12. And with a barrier of $-2.0$ ?
13. Why does a higher barrier (earlier delisting) raise the bias?
14. What would a survivors-only backtest of a strategy that buys stocks after large falls report?

**Part IV — Judgement.**

15. Where do you get the delisted names for a real backtest?
16. What delisting return do you use when a database has none, and why not zero?
17. How does the bias interact with an equal-weighted versus a capitalisation-weighted universe?
18. State the *named result* : the annual return inflation of the survivors-only universe at the simulated delisting rate, and the share of it that the delisting month explains.
19. Name two other biases with the same structure.
20. In one sentence: why is a list of today’s companies a forecast?

**Solution of Problem 3.1.**

**1.** $1 - e^{-1.6} = 79.8\%$. **2.** 2.2% of names a year. **3.** 7.7% a year. **4.** 12.9%. **5.** 5.2 points. **6.** 20.9%. **7.** $r_F = -11.1\%$, $r_S = 12.6\%$ a year. **8.** $0.209 \times (12.6 +
11.1) = 5.0$ points; Question 5 averages month by month over a changing number of survivors (few early, many late) instead of pooling name-months. **9.** 4.5 points. **10.** 6.7 points. **11.** 3.5% a year, 7.0 points. **12.** 1.4% a year, 3.8 points. **13.** More names reach the barrier, so more name-months belong to failures ($w$ rises), while each failure’s decline is still steep. **14.** A strong profit: the fallen stocks that went on to fail are missing, and those that remain are by construction the ones that recovered; the true result may have the opposite sign. **15.** From a database that keeps delisted securities and their delisting returns, with universes as of each date. **16.** An estimate by delisting reason, such as the $-30\%$ Shumway recovered for performance delistings; zero is optimistic, since performance delistings are surprises that the last traded price does not reflect. **17.** Failing names are small when they fail, so a capitalisation-weighted universe gives them little weight and suffers less; equal weighting, common in research, suffers most. **18.** *Named result:* the survivors-only universe overstates the equal-weighted return by 5.2 points a year when 2.2% of names delist each year, and the delisting month explains 0.7 of them (the bias is still 4.5 points with a delisting return of zero). **19.** [Backfill bias](#def-rs-point-in-time-data-and-the-biases-vintage) in fund and vendor databases, and testing on today’s index members. **20.** It lists the companies that will survive, which no one knew at the start of the sample.

## 3.10 Interview questions

**Interview question 3.1 ★ researcher, trader.**

What is [survivorship bias](#def-rs-point-in-time-data-and-the-biases-survivorship), and how would it show up in a stock-selection backtest?

**Solution of Interview question 3.1.**

Studying only securities that exist at the end of the sample. In a stock-selection backtest built on today’s listed companies, the failures are missing, so strategies that buy cheap, distressed or small stocks look far better than they were, and average returns are inflated by several points a year in equal-weighted universes.

*What the interviewer is looking for: that the missing names are the losers, and the fix: a universe as of each date with delistings.*

**Interview question 3.2 ★★ researcher.**

Your macro model uses GDP growth. Which GDP figure should it use for each date of the backtest?

**Solution of Interview question 3.2.**

The estimate available at each decision date: the vintage published by then (first or later release), not today’s revised figure; the model must also be fitted on those real-time values, since it will be fed them live.

*What the interviewer is looking for: real-time vintages; the distinction between valid and [knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal).*

**Interview question 3.3 ★★ developer, researcher.**

Design a table for fundamental data that supports point-in-time queries. What are its keys?

**Solution of Interview question 3.3.**

Key (company, field, fiscal period, [knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal)) with the value, where [knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal) is the filing or vendor delivery time; never overwrite. A query takes (company, field, period, as-of time) and returns the latest row known at that time; a snapshot query returns every period as known then.

*What the interviewer is looking for: two time columns, append-only storage, as-of semantics.*

**Interview question 3.4 ★★ developer, mle.**

How would you test automatically that a feature pipeline has no look-ahead?

**Solution of Interview question 3.4.**

Perturbation: for each decision time, replace every input value with a [knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal) after it by noise and check that the features at that decision are unchanged. Also a guard in the data access layer that raises on any later [knowledge time](#def-rs-point-in-time-data-and-the-biases-bitemporal), and a comparison of the backtest run on stored vintages with the same run on today’s data (a large difference flags a sensitivity to revisions).

*What the interviewer is looking for: an automated test, not a code review.*

**Interview question 3.5 ★★ researcher, risk.**

A hedge-fund index shows high returns and low volatility. Which database biases would you suspect?

**Solution of Interview question 3.5.**

Survivorship (dead funds dropped), backfill (funds join with their good past), self-selection (funds report when they do well and stop when they do badly), and stale or smoothed marks for illiquid holdings, which lower measured volatility.

*What the interviewer is looking for: several distinct biases and their direction.*

**Interview question 3.6 ★★★ researcher.**

Show that a survivors-only average overstates a universe’s average return by $w(r_S - r_F)$, and say which of the terms is larger in practice.

**Solution of Interview question 3.6.**

Pooled over name-months, the universe average is $(1 - w)r_S + wr_F$; the survivors’ average is $r_S$; the difference is $w(r_S
- r_F)$. In practice the gap $r_S - r_F$ is dominated by the months failing names spend declining, not the delisting month: in the chapter’s simulation, 4.5 of the 5.2 points remain with a delisting return of zero.

*What the interviewer is looking for: the identity, and where the gap comes from.*
