---
title: "Fundamental, Analyst and Event Features"
book: "Research Craft: Predictors, Backtests, Measurement, Portfolios"
subject: quant
language: en
chapter: 11
exercises: 8
source: https://one-course.com/books/quant/7/en/chapter/11-fundamental-analyst-and-event-features
---

# Chapter 11 — Fundamental, Analyst and Event Features

A company closes its quarter on 31 December, announces its earnings on 25 January and files its quarterly report on 5 February. A database keyed on the period’s end shows the quarter’s earnings on 31 December, and a backtest built on it trades the announcement’s news three and a half weeks before anyone could read it. The error is invisible in the code, which joins on a date column like any other, and spectacular in the results. This chapter keys accounting numbers, surprises, analyst forecasts and event dates on the time they became known. Its real data are ten large US companies’ reported earnings and filing calendars from the SEC’s public EDGAR interfaces; its measurement of the damage is on `firm.synthmkt`, where the truth is known.

## 11.1 Accounting data, point-in-time

**Definition 11.1 (Fiscal period, filing lag).**

A *fiscal period* is the quarter or year an accounting report describes, identified by its start and end dates; the [valid time](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-bitemporal) of every number in the report is the period. The *filing lag* is the time from the period’s end to the filing of the report with the regulator.

An accounting number has at least three dates after its period ends: the earnings release (a press release, filed in the US as a Form 8-K under item 2.02), the periodic report (Form 10-Q for a quarter, 10-K for a year), and any later filing that reports the same period again as a comparative, possibly with a different value ([Figure 11.1](#fig-rs-fundamental-analyst-and-event-features-timeline)). The market learns the number at the release; a database built from the periodic reports learns it at the filing; a database that stores only the period end pretends to know it at the end.

![The dates of a quarterly number, at the medians of ten large US companies’ first three fiscal quarters since 2013 (23 days to the earnings release, 29 to the 10-Q), with the SEC’s filing deadlines.](https://one-course.com/images/onecourse/chapters/quant-7/rs-fundamental-analyst-and-event-features/fig-381d4e5f940f.svg)

***Figure 11.1.** The dates of a quarterly number, at the medians of ten large US companies’ first three fiscal quarters since 2013 (23 days to the earnings release, 29 to the 10-Q), with the SEC’s filing deadlines.*

**As of September 2026 — US periodic report deadlines.**

The SEC’s Form 10-Q instructions require the quarterly report within 40 days of the quarter’s end for large accelerated and accelerated filers and 45 days for all other registrants; no 10-Q is filed for the fourth quarter. Form 10-K is due 60 days after the fiscal year’s end for large accelerated filers, 75 for accelerated filers and 90 for all others. The SEC’s EDGAR application programming interfaces serve every filer’s submission history and the XBRL data of its financial statements as JSON, updated in real time, without authentication or keys.

The ten companies of `rs_fetch_edgar.py` (Apple, Microsoft, Walmart, Coca-Cola, Johnson & Johnson, Exxon Mobil, Procter & Gamble, Intel, Home Depot and Nike) have 542 [fiscal period](#def-rs-fundamental-analyst-and-event-features-fiscal) ends since 2013 in their EDGAR submission histories, 541 of them followed by an earnings release within 90 days. For the first three quarters the median company released its earnings 23 days after the period ended and filed its 10-Q 29 days after; for fiscal years the medians are 26 and 48 days. The largest 10-Q lag is 41 days and the largest 10-K lag 60, inside the deadlines. Only 16% of reports were filed on the day of the release; the median gap is a week, during which the number is public but absent from a database fed by filings ([Figure 11.2](#fig-rs-fundamental-analyst-and-event-features-calendar)).

![Median days from the end of each of the first three fiscal quarters to the earnings release and to the 10-Q filing, ten companies, 2013 to 2026; the dashed line is the 40-day deadline. Some file with the release, some weeks later. Data: SEC EDGAR submissions, via rs_fetch_edgar.py.](https://one-course.com/images/onecourse/chapters/quant-7/rs-fundamental-analyst-and-event-features/fig-b788fccc2eed.svg)

***Figure 11.2.** Median days from the end of each of the first three fiscal quarters to the earnings release and to the 10-Q filing, ten companies, 2013 to 2026; the dashed line is the 40-day deadline. Some file with the release, some weeks later. Data: SEC EDGAR submissions, via `rs_fetch_edgar.py`.*

The numbers also change after they are first filed. Each XBRL fact in EDGAR lists every filing that reported it, so a quarter’s diluted earnings per share can be followed from its first report through every later comparative. Among the ten companies, 31 quarterly values changed:

| company | values changed | by an exact ratio | ratios |
| --- | --- | --- | --- |
| Apple | 13 | 13 | 7 and 4 |
| Nike | 7 | 6 | 2 |
| Microsoft | 5 | 0 | — |
| Walmart | 3 | 3 | 3 |
| Coca-Cola | 2 | 2 | 2 |
| Johnson & Johnson | 1 | 0 | — |
| the other four | 0 | 0 | — |

Twenty-four of the 31 changed by an exact ratio (to the cent), the signature of a share split: Apple’s 10-Ks record a seven-for-one split on 6 June 2014 and a four-for-one split on 28 August 2020, after which all per-share information was retroactively adjusted. The other seven are [restatements](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-vintage) of substance. Microsoft’s diluted earnings for its fiscal year 2017 were first reported as $2.71 and a year later as $3.25: it adopted a new revenue standard in its fiscal 2018 using the full retrospective method, which required it to restate each prior period presented. A backtest of fiscal 2017 must use $2.71 until August 2018.

**Definition 11.2 (Trailing twelve months).**

The *trailing twelve months* (TTM) value of a flow (earnings, revenue, cash flow) at a quarter’s end is the sum of the four most recent consecutive quarters, each as known at the decision time; a fourth quarter that is reported only inside an annual figure is the year minus the first nine months.

Summing four quarters is where the dates bite together. Apple’s first-reported quarterly diluted EPS for its fiscal 2014 were $14.50 and $11.62 (before the June split) and $1.28 and $1.42 (after): a TTM built from first reports is $28.82, while the annual figure Apple reported for the same year was $6.45. The same happens across the 2020 split, $10.85 against an annual $3.28 ([Figure 11.3](#fig-rs-fundamental-analyst-and-event-features-ttm)). Today’s view has the opposite problem: it rescales history with knowledge of later splits, and it is itself a patchwork, because a period is rescaled only if a later filing reported it again. The rule is to store per-share values as first reported with their [knowledge times](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-bitemporal), store splits as events with theirs (chapter 4), and rescale at query time to the share basis of the decision date, the basis of the price the number will be divided by.

![Apple’s trailing-twelve-month diluted EPS at each quarter’s end, summed from each quarter’s first-reported value and from its latest value in EDGAR, with the two stock splits. Neither is a usable series without split adjustment at the decision date. Data: SEC EDGAR company facts.](https://one-course.com/images/onecourse/chapters/quant-7/rs-fundamental-analyst-and-event-features/fig-a7d6baf7ec3a.svg)

***Figure 11.3.** Apple’s trailing-twelve-month diluted EPS at each quarter’s end, summed from each quarter’s first-reported value and from its latest value in EDGAR, with the two stock splits. Neither is a usable series without split adjustment at the decision date. Data: SEC EDGAR company facts.*

## 11.2 Ratios as features

**Definition 11.3 (Accounting ratio).**

An *accounting ratio* divides one accounting quantity by another or by a market quantity: book-to-price, earnings yield (TTM earnings over price), sales-to-price, return on equity, leverage, asset growth.

A ratio has two [knowledge times](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-bitemporal) and must respect both. The accounting side is known at its filing (or release); the market side (price, market capitalisation) is known at its own time, normally the decision time. The pairings that fail are regular: a TTM earnings figure divided by a price on a different share basis; a book value keyed on its period end; a ratio whose denominator can change sign (earnings, book equity), which needs a separate treatment for negative values rather than a rank that puts the most negative earnings yield next to the most expensive stock. Book-to-price in `firm.synthmkt` shows how much the key date matters for a slow variable: its mean rank IC over the next 21 days is 0.029 (standard error 0.038) keyed on the period end and $-0.005$ (0.023) keyed on the filing day. The difference is noise: each quarterly IC is dominated by the value factor’s realised return over its window, and moving the window a few weeks changes little in expectation because nothing in the simulated market happens on the filing day. The damage done by a wrong key date is proportional to how much the number’s release moves the price, and a book value’s release does not move it much.

## 11.3 Surprises and revisions

**Definition 11.4 (Earnings surprise, standardised unexpected earnings).**

An *earnings surprise* is the reported earnings of a period minus their expected value: an analysts’ consensus forecast (Book 2, chapter 31) or a time-series model. *Standardised unexpected earnings* (SUE) divide the surprise of a seasonal random walk, $x_q - x_{q-4}$, by the standard deviation of that difference over the previous eight quarters.

Foster, Olsen and Shevlin (1984) documented that systematic post-announcement drifts in returns are associated with the sign and magnitude of unexpected earnings, and that for expectation models based on the time series of reported quarterly earnings the forecast error and firm size explained 81% and 61% of the variation in the drifts. The release is the event; the drift after it is the feature’s opportunity; a key date before the release captures the release itself. In `firm.synthmkt` each quarter’s earnings are announced a random number of trading days after its end (31 on average, about six weeks: the head start of a period-end key), with a jump in the direction of the true surprise and a smaller drift afterwards. The surprise keyed on its announcement has a mean rank IC with the next 21 days’ market-adjusted return of 0.020 (standard error 0.006): the planted drift. Keyed on the period end it is 0.167, eight times larger, because a third of the windows contain the announcement’s jump ([Figure 11.4](#fig-rs-fundamental-analyst-and-event-features-ic)).

![Mean rank IC over 39 quarterly cross-sections of firm.synthmkt (bars: two standard errors) of the true earnings surprise and of book-to-price, keyed on the day each became known (announcement, filing) and on the period’s end. The period-end key multiplies the surprise’s IC by eight and leaves book-to-price within its noise.](https://one-course.com/images/onecourse/chapters/quant-7/rs-fundamental-analyst-and-event-features/fig-d0af86192f33.svg)

***Figure 11.4.** Mean rank IC over 39 quarterly cross-sections of `firm.synthmkt` ([bars](https://one-course.com/books/quant/7/en/chapter/2-market-data-for-research#def-rs-market-data-for-research-bar): two standard errors) of the true [earnings surprise](#def-rs-fundamental-analyst-and-event-features-sue) and of book-to-price, keyed on the day each became known (announcement, filing) and on the period’s end. The period-end key multiplies the surprise’s IC by eight and leaves book-to-price within its noise.*

**Definition 11.5 (Analyst revision, forecast dispersion).**

An *analyst revision* is a change in an analyst’s forecast or recommendation for a security; as a feature, the change in the consensus forecast over a window, scaled by price or by the forecasts’ dispersion. *Forecast dispersion* is the cross-analyst standard deviation of the current forecasts divided by the absolute value of their mean.

Womack (1996) found that new buy and sell recommendations by analysts at major US brokerages move prices at once, although few coincide with new public news, and that the reactions are incomplete: the mean drift after a buy recommendation was modest (+2.4%) and short-lived, after a sell recommendation larger ($-9.1\%$) and six months long. Diether, Malloy and Scherbina (2002) found that stocks with higher dispersion of analysts’ earnings forecasts earn lower future returns than otherwise similar stocks, most so in small stocks and past losers, consistent with prices reflecting the optimists’ view when the most pessimistic investors do not trade. Both are point-in-time problems first: a forecast database must keep each forecast with the date the analyst made it and the date the vendor recorded it, and a recommendation that is later withdrawn or corrected must stay in history as it was. Coverage is its own bias: analysts cover the firms they expect to do well and drop the others, so the set of covered firms changes with the outcome.

## 11.4 Event calendars

**Definition 11.6 (Event calendar, announcement window).**

An *event calendar* lists scheduled corporate events (earnings releases, dividend dates, shareholder meetings, index changes) with the date each was announced and the date it occurs. An *announcement window* is a fixed range of trading days around an event, such as $[-1, +1]$, used to measure its reaction or to exclude it.

A calendar is itself [point-in-time data](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-lookahead): a scheduled date is announced before the event, may move, and is stored with the time it was announced; a date that moves is itself worth testing as a feature. The ten companies’ release lags are regular: Home Depot released 16 days after the end of every one of its first three fiscal quarters since 2013, Exxon Mobil 30 days give or take two, while Apple’s varies more (a standard deviation of five days). A predicted date is therefore a usable feature before the confirmed date exists, and *days to the next event* is a conditioning variable every research system needs: it switches features on (a surprise after the release), off (a volatility estimate across a release), and sizes risk around the known unknown. [Announcement windows](#def-rs-fundamental-analyst-and-event-features-calendar) serve the other way round: a reversal or volatility feature estimated through an earnings day measures the announcement, not the effect it was meant to measure, so research code excludes the window or models it separately.

## 11.5 Predictor cards

**Predictor card 11.1 — Standardised unexpected earnings.**

**Definition.** $(x_q - x_{q-4})/\sigma_8$, the seasonal random-walk surprise of quarterly EPS over the standard deviation of the last eight such differences; or the surprise against the consensus, over the dispersion.

**Inputs and timestamps.** Quarterly EPS as first released, keyed on the release (the 8-K) or, from filings only, on the filing; split factors as known at the decision date.

**Rationale.** Post-announcement drift (Foster, Olsen and Shevlin, 1984).

**Horizon and half-life.** Weeks to a quarter after the release; on `firm.synthmkt`, IC 0.020 over 21 days with the correct key.

**Normalisation.** Rank across the stocks that announced in the window; neutralise size.

**Failure modes.** A period-end key (IC 0.167 on the simulated market: all of it the announcement); a surprise measured against a restated number; a stale consensus.

**Sources.** As cited; `firm.fundpit.sue`; `rs_fundamentals.ic_inflation`.

**Predictor card 11.2 — Book-to-price, as known.**

**Definition.** The latest book equity known at the decision time, over the market capitalisation at the decision time.

**Inputs and timestamps.** Book equity keyed on its filing; shares and price at the decision date, on one share basis.

**Rationale.** The value premium (Book 8, chapter 6).

**Horizon and half-life.** Months to years; the book changes quarterly, the price daily.

**Normalisation.** Rank within industry; negative book equity handled apart.

**Failure modes.** A period-end key (little damage for a slow ratio: on the simulated market the difference is within noise); mismatched share bases; restated book values used before their [restatement](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-vintage).

**Sources.** `rs_fundamentals.ic_inflation`.

## 11.6 Tutorial: the five-week head start

**Goal.** Build the dates of real earnings numbers from EDGAR, store them point-in-time, derive quarters and TTM values as known at any date, and measure on a synthetic market what a period-end key does to a surprise and a ratio. **End state:** Figures [11.2](#fig-rs-fundamental-analyst-and-event-features-calendar), [11.3](#fig-rs-fundamental-analyst-and-event-features-ttm) and [11.4](#fig-rs-fundamental-analyst-and-event-features-ic); the [restatement](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-vintage) table.

1. **Fetch** with `rs_fetch_edgar.py` : submissions (8-K item 2.02, 10-Q, 10-K) and company facts (diluted EPS with every filing that reported it), with a descriptive User-Agent and a pause between requests.
2. **Store** each fact under its duration, valid at the period end, known at the filing. `def duration_kind (start, end) -> str | None : n = (_d(end) - _d(start)).days return next ((k for k, lo, hi in KINDS if lo <= n <= hi), None ) def load (facts, field: str = " eps " ) -> Store: """facts: iterable of dicts with entity, start, end, val, filed.""" s = Store() for f in facts: k = duration_kind(f[" start " ], f[" end " ]) if k is not None : s.put(f[" entity " ], f " { field} _ { k} " , f[" end " ], f[" filed " ], float (f[" val " ])) return s` **Listing 11.1.** Facts into a point-in-time store, by duration. code/firm/fundpit/firm_fundpit.py
3. **Quarters as known**, the fourth derived from the year and the nine months. `def quarterly (store: Store, entity, field: str , known) -> dict : """Quarters as known at `known`; a fourth quarter is the year minus the nine months ending one quarter earlier (or minus the three quarters) when those are known too.""" q = dict (store.snapshot(entity, f " { field} _Q " , known)) nine = store.snapshot(entity, f " { field} _9M " , known) for end, fy in store.snapshot(entity, f " { field} _FY " , known).items(): if end in q: continue e = _d(end) near = [v for k, v in nine.items() if 80 <= (e - _d(k)).days <= 100 ] prev = sorted (k for k in q if 0 < (e - _d(k)).days <= 290 ) if near: q[end] = fy - near[0 ] elif len (prev) == 3 : q[end] = fy - sum (q[k] for k in prev) return dict (sorted (q.items()))` **Listing 11.2.** Quarterly values as known at a date. code/firm/fundpit/firm_fundpit.py
4. **Run** `calendar_summary()` , `restatements()` , `apple_ttm()` , `ic_inflation()` and `fig_fundamentals.py` .

**What to change next.** Key the surprise on the 10-Q filing instead of the release and measure what the week between them costs; rescale Apple’s first-reported quarters to the share basis of each decision date and rebuild a clean TTM series.

## 11.7 Build: the fundamental feature store

**Purpose.** Accounting features for the miniature firm’s research that cannot leak: every value carries its period and its [knowledge time](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-bitemporal), and every derived value (quarters, TTM, SUE, ratios) is computed from what was known at the decision date.

**Interface.** `duration_kind`, `load(facts, field)`, `quarterly(store, entity, field, known)`, `ttm`, `sue(values, n)`, `restated`, `split_ratio`, `revision`, `dispersion`, `days_to_event`, `in_window`; storage is `firm.pit.Store`.

**Rules.** [Valid time](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-bitemporal) is the period end, [knowledge time](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-bitemporal) the filing (or the release when it is stored); a fourth quarter exists only when the year and the nine months are both known; a per-share value is never compared across share bases.

**Acceptance tests.** `code/firm/fundpit/tests/`: durations of a hand-built year, the fourth quarter ($4.80 minus $3.30), a quarter unknown before its filing, TTM, the later view after a split; [restatement](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-vintage) listing and split ratios to the cent; SUE of a seasonal jump and of a random walk; revision, dispersion, days to the next event and windows by hand.

**Stretch.** Release dates from 8-K filings as the [knowledge time](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-bitemporal) of each quarter; a consensus store with analyst and vendor timestamps; the split events of `firm.secmaster` applied at query time.

Sources and further reading

- US Securities and Exchange Commission, Form 10-Q and Form 10-K, general instructions; EDGAR application programming interfaces (data.sec.gov).
- Apple Inc., Form 10-K for fiscal 2014 and fiscal 2020; Microsoft Corporation, Form 10-K for fiscal 2018 (SEC EDGAR).
- G. Foster, C. Olsen and T. Shevlin, “Earnings releases, anomalies, and the behavior of security returns”, *The Accounting Review* 59(4), 1984.
- K. L. Womack, “Do brokerage analysts’ recommendations have investment value?”, *Journal of Finance* 51(1), 1996.
- K. B. Diether, C. J. Malloy and A. Scherbina, “Differences of opinion and the cross section of stock returns”, *Journal of Finance* 57(5), 2002.

## 11.8 Exercises

**Exercise 11.1 ★.**

A company’s fiscal year ends on 30 June. It reports nine-month diluted EPS of $3.30 at 31 March and annual EPS of $4.80. What is the fourth quarter’s EPS, and when is it first known if the 10-K is filed 55 days after the year’s end?

**Solution of Exercise 11.1.**

$\$4.80 - \$3.30 = \$1.50$. It is first known when the 10-K is filed, 55 days after 30 June: 24 August (or at the earnings release, if the store keys on releases).

**Exercise 11.2 ★.**

An accelerated filer’s quarter ends on 30 September. By which date must its 10-Q be filed? And a large accelerated filer’s 10-K for a year ending 31 December (not a leap year)?

**Solution of Exercise 11.2.**

Forty days after 30 September: 9 November. Sixty days after 31 December in a year of 365 days: 1 March. (A deadline that falls on a weekend or holiday moves to the next business day.)

**Exercise 11.3 ★.**

A quarter’s EPS was first reported as $1.42 and later as $0.355. What happened, and how should the store keep both?

**Solution of Exercise 11.3.**

$1.42/0.355 = 4$: a four-for-one split rescaled the history. The store keeps $1.42 with its first filing date as [knowledge time](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-bitemporal) and $0.355 with the later filing’s, and records the split as an event with its own date; a query rescales to the share basis of its decision date.

**Exercise 11.4 ★★.**

Quarterly EPS over twelve quarters are 1.00, 1.10, 0.90, 1.20, 1.04, 1.12, 0.96, 1.22, 1.06, 1.16, 0.94, 1.40. Compute the seasonal differences and the SUE of the last quarter with $n = 6$ (the six previous differences, sample standard deviation).

**Solution of Exercise 11.4.**

The seasonal differences are 0.04, 0.02, 0.06, 0.02, 0.02, 0.04, $-0.02$ and 0.18. The six before the last (0.02, 0.06, 0.02, 0.02, 0.04, $-0.02$) have a sample standard deviation of 0.0266, so the last quarter’s SUE is $0.18/0.0266 = 6.8$.

**Exercise 11.5 ★★.**

If announcements fall uniformly between 0 and 62 trading days after a quarter’s end, what share of 21-day windows starting the day after the quarter’s end contain the announcement? Relate it to the eightfold IC inflation.

**Solution of Exercise 11.5.**

The window from day 1 to day 21 after the quarter’s end contains the announcement when it falls on one of those 21 of the 63 possible days: a third (0.330 measured). In those windows the feature sees the announcement’s jump, whose size dwarfs the drift the correct key measures; a third of the windows carrying a jump several times larger than a quarter’s drift is enough to multiply the IC by eight.

**Exercise 11.6 ★★.**

Why can book-to-price keyed on the period end show little inflation while the [earnings surprise](#def-rs-fundamental-analyst-and-event-features-sue) shows a lot? What would make book-to-price inflate too?

**Solution of Exercise 11.6.**

A look-ahead gains what the price does between the period end and the [knowledge time](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-bitemporal) because of the hidden number. An [earnings surprise](#def-rs-fundamental-analyst-and-event-features-sue) moves the price on its release, inside the gap; a book value’s filing moves it little, and book-to-price changes slowly anyway, so the window’s return is nearly independent of which of the two keys is used. Book-to-price would inflate if the book value carried news the price reacted to on the filing (a large write-down), or if the period-end price used in the ratio were replaced by a later one (a price look-ahead).

**Exercise 11.7 ★★★.**

*Coding.* From `data/research/edgar_filings.csv`, compute for each company the share of quarters in which the 10-Q was filed on the release day, and the median days from release to filing for the others.

**Solution of Exercise 11.7.**

`rs_fundamentals.release_to_filing`: Microsoft filed every 10-Q on the release day, Procter & Gamble 95% of them; Coca-Cola 7% and Intel 10%. For the others, the median days from release to filing are one (Apple, Intel), two (Coca-Cola, Procter & Gamble), five (Exxon Mobil), seven (Home Depot), 14 (Johnson & Johnson, Nike) and 16 (Walmart).

**Exercise 11.8 ★★★.**

*Find the flaw.* “We use the vendor’s historical fundamentals table, which is updated whenever a company restates, so our earnings history is always the most accurate available.”

**Solution of Exercise 11.8.**

It is the most accurate history and the wrong one for a backtest: every value is today’s, so a decision in the past sees [restatements](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-vintage) (Microsoft’s fiscal 2017 EPS of $3.25 instead of the $2.71 known until August 2018) and split rescalings it could not have known. Research needs the values as first reported, with every later vintage and its date.

## 11.9 Problem: The Five-Week Head Start

**Problem 11.1.**

Weekend problem — a key date, measured

Ten companies’ EDGAR data since 2013, and `firm.synthmkt` with its earnings announcements, filings and book values.

**Part I — The dates.**

1. How many period ends have the ten companies since 2013, and how many are followed by a release within 90 days?
2. What are the median days from a quarter’s end to the release and to the 10-Q, and for a fiscal year to the release and the 10-K?
3. What are the largest 10-Q and 10-K lags, and how do they compare with the deadlines?
4. How often is the report filed on the release day, and what is the median gap?

**Part II — The values.**

5. How many quarterly values changed after their first filing, and how many by an exact ratio?
6. What were Apple’s two splits, and what did they do to its per-share history?
7. What happened to Microsoft’s fiscal 2017 EPS, and why?
8. What TTM does Apple’s fiscal 2014 give from first-reported quarters, and what annual EPS did Apple report?
9. How should a store keep split-affected values?

**Part III — The damage.**

10. What is the average head start of a period-end key in `firm.synthmkt` ?
11. What are the surprise’s IC keyed on the announcement and on the period end, with their standard errors?
12. What share of the windows contain the announcement?
13. What are book-to-price’s ICs under the two keys, with standard errors?
14. Why is the book-to-price difference not significant?
15. State the *named result* : the IC inflation of a surprise and a value signal when period-end dating replaces known-date dating.

**Part IV — Beyond.**

16. Key the surprise on the filing instead of the release: what do you expect to lose?
17. Which is worse for a backtest, a period-end key or a latest-value (restated) history, for a SUE feature?
18. How would you detect a period-end key in someone else’s backtest?
19. What should a fundamentals vendor deliver for research use?
20. In one sentence: what is the [knowledge time](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-bitemporal) of an accounting number?

**Solution of Problem 11.1.**

1. 542 period ends; 541 are followed by a release within 90 days.
2. Quarters: 23 days to the release, 29 to the 10-Q. Fiscal years: 26 to the release, 48 to the 10-K.
3. 41 and 60 days; the deadlines are 40 (moved to the next business day when it falls on a holiday or weekend) and 60 days for large filers.
4. 16% on the same day; the median gap is 7 days.
5. 31, of which 24 by an exact ratio.
6. Seven-for-one on 6 June 2014 and four-for-one on 28 August 2020; per-share values before each were later reported divided by 7 and by 4 in the filings that repeated them.
7. First reported at $2.71, restated to $3.25 in the fiscal 2018 filings: Microsoft adopted a new revenue standard with the full retrospective method, which restated each prior period presented.
8. $28.82 from first-reported quarters (two before the split, two after), against an annual EPS of $6.45.
9. As first reported with their [knowledge times](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-bitemporal) , with splits stored as events, and rescaled at query time to the decision date’s share basis.
10. 31 trading days, about six weeks.
11. 0.020 (standard error 0.006) keyed on the announcement; 0.167 (0.006) keyed on the period end.
12. A third (0.330).
13. 0.029 (0.038) keyed on the period end; $-0.005$ (0.023) keyed on the filing, with the period-end price.
14. Each quarter’s IC is dominated by the value factor’s return in its window, and nothing happens on the filing day to make the two windows’ returns depend on the ratio differently: the expected inflation is small and the noise large.
15. **Named result.** In `firm.synthmkt` , keying the [earnings surprise](#def-rs-fundamental-analyst-and-event-features-sue) on the period end instead of its announcement multiplies its IC by eight (0.020 to 0.167); keying book-to-price on the period end instead of its filing changes it by less than its standard error.
16. The days between the release and the filing (a week at the median in the EDGAR data): for a drift feature, the start of the drift.
17. The period-end key, by far, for a SUE feature: it trades the announcement itself. A restated history distorts the size of some surprises but not their timing.
18. Recompute the feature’s IC with the entry shifted to the [knowledge time](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-bitemporal) ; plot the IC against the shift, since a leak shows as a cliff at the true publication lag; check the timing of the returns earned around announcements.
19. Every value with its period, its first publication time, every later vintage with its time, and corporate actions as dated events.
20. The time it was first published, at the release or, for a database fed by filings, at the filing, never the end of the period it describes.

## 11.10 Interview questions

**Interview question 11.1 ★ researcher.**

What is point-in-time fundamental data, and why does it matter?

**Solution of Interview question 11.1.**

Data stored with the time each value became known, and queried as of a decision time, so a backtest sees only what was known then. It matters because accounting numbers are published weeks after their periods and are later restated or rescaled; a database of current values keyed on period ends leaks both.

**Interview question 11.2 ★★ researcher.**

How would you compute trailing-twelve-month earnings for a company that reports only year-to-date figures?

**Solution of Interview question 11.2.**

Difference the year-to-date figures: the second quarter is the six months minus the first, the fourth the year minus the nine months; each derived quarter is known when its later component is. Sum four consecutive quarters as known at the decision date, on one share basis for per-share figures.

**Interview question 11.3 ★★ researcher, trader.**

What is post-earnings-announcement drift, and how would you build a feature for it?

**Solution of Interview question 11.3.**

The tendency of returns after an earnings release to continue in the direction of the surprise (Foster, Olsen and Shevlin, 1984). Feature: SUE against a seasonal random walk or the consensus, keyed on the release, ranked among the stocks that announced recently, held for weeks; check the key date and neutralise size.

**Interview question 11.4 ★★ researcher.**

Your value signal’s backtest improves by a third when you switch from the filing date to the period end as the key. Which is right, and why might the improvement be small or large?

**Solution of Interview question 11.4.**

The filing date is right: the period end uses numbers before they were known. The improvement is large when the number’s publication moves the price (earnings) and small when it does not (a slowly changing book value); on the chapter’s simulated market book-to-price moved by less than its standard error, the surprise by a factor of eight. A third is a warning either way.

**Interview question 11.5 ★★ researcher.**

What biases affect analyst forecast data?

**Solution of Interview question 11.5.**

Timestamps (forecast date against the date the vendor recorded it), revisions that overwrite history, survivorship of analysts and of covered firms, coverage chosen on expected performance, stale forecasts in the consensus, and optimism that varies with the analyst’s incentives.

**Interview question 11.6 ★★★ researcher, developer.**

Design the storage of per-share accounting values so that any query returns what was known at a date, on the share basis of that date.

**Solution of Interview question 11.6.**

A [bitemporal](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-bitemporal) store of facts keyed by (entity, field, period) with every vintage and its [knowledge time](https://one-course.com/books/quant/7/en/chapter/3-point-in-time-data-and-the-biases#def-rs-point-in-time-data-and-the-biases-bitemporal), values kept as reported on their own share basis; a separate table of split events with their dates; a query (entity, field, period, decision time) that takes the last vintage known then and multiplies it by the product of the split factors between the vintage’s basis and the decision date’s, both taken as known at the decision time.
