---
title: "Text: From Bag of Words to Embeddings"
book: "Machine Learning for Markets"
subject: quant
language: en
chapter: 13
exercises: 8
source: https://one-course.com/books/quant/12/en/chapter/13-text-from-bag-of-words-to-embeddings
---

# Chapter 13 — Text: From Bag of Words to Embeddings

A headline says a company beat estimates, and its shares fall. A dictionary scores the headline positive; a model trained on returns learns that “beat estimates but cut guidance” is bad news; and neither knows which of three companies called Apex the headline is about. Book 7 (chapter 12) introduced text as data and its first instrument, the dictionary; Book 8 (chapter 17) traded machine-readable news. This chapter builds the machinery between the words and the forecast: how a document becomes a vector, how the vector is scored against returns, what topics and embeddings can and cannot see, and how a mention becomes a security. On a synthetic corpus with a known truth, a return-supervised model reaches an information coefficient of 0.27 where a dictionary reaches 0.22, embeddings learned without returns reach 0.09 because they put “beat” next to “missed”, and an entity linker that uses today’s tickers maps one mention in eight to the wrong company or to none.

## 13.1 From documents to vectors

**Definition 13.1 (Tokenisation, bag of words, n-gram, tf-idf).**

*Tokenisation* splits a text into units (words, word pieces, punctuation), usually after normalising case. A *bag of words* represents a document by the counts of its tokens, forgetting their order. An *n-gram* is a sequence of $n$ consecutive tokens; bags of n-grams keep some order. *Term frequency–inverse document frequency* (tf-idf) weights the count of term $w$ in document $d$ by $\log(N/n_w)$, where $n_w$ of the $N$ documents contain $w$, so that words found everywhere count for little; the rows are then scaled to unit length.

The chapter’s corpus has 20 000 headlines over six years (1 512 trading days) about 32 companies: 28 at the start, three of them called Apex (a software firm, a miner and a healthcare firm), six ticker changes, three renames, and four delistings whose tickers are reused by new listings 30 days later. Each headline names a company (by its full name, its short name or its ticker), states one of 19 kinds of event in one of several phrasings, and sometimes mentions a sector word; each carries a reaction return equal to the event’s planted effect plus noise of 3% standard deviation. The first 60% of the days (11 923 headlines) train; the last 40% (8 077) test.

A headline such as “Apex beat estimates but cut guidance amid copper demand” passes through three steps before any model sees it. The company mention is recognised and replaced by a placeholder, so that no model learns a company’s name as a signal; the text is lowercased and split into the tokens `<co> beat estimates but cut guidance amid copper demand`; and the bigrams add `<co> beat`, `beat estimates`, `estimates but`, `but cut`, `cut guidance` and three more. The training headlines’ 123 words (those found in at least three headlines) become 473 features with bigrams; in real news the vocabulary runs to tens of thousands and bigrams multiply it again.

## 13.2 Sentiment from dictionaries and from returns

A dictionary (Book 7, chapter 12) is a fixed instrument: the chapter’s hand-built lists hold 11 positive and 16 negative words. It reads a result correctly (“beat” scores $+0.20$ on average, against a planted effect of $+1\%$; a guidance cut $-0.18$ for $-1.5\%$) and misreads everything built from more than one word: “beat estimates but cut guidance” scores exactly zero, one hit each way, where the planted effect is $-1.2\%$; “did not cut guidance” scores negative for a planted $+0.3\%$; “wins industry award” scores positive for nothing; “announces share buyback” scores nothing for $+0.7\%$. Loughran and McDonald (2011) showed that general word lists misclassify financial words; any list misses the words its authors did not think of and cannot weigh the ones it has.

The alternative is to let returns choose the words. The return-supervised screen of [Listing 13.1](#lst-ml-text-supervised) keeps the words whose headlines’ mean return differs from the others’ by more than three standard errors and weighs each by that difference, in the spirit of Ke, Kelly and Xiu (2019), whose screening-and-topic model builds a sentiment score adapted to return prediction. On the training half it keeps 41 words, from “trims” ($-1.43\%$) and “halves” to “takeover” ($+1.79\%$); it also keeps “to” and “be”, which carry information only because of how the corpus’s takeover headlines are phrased, the kind of word a stop list would have removed and a return screen rightly keeps. [Table 13.1](#tab-ml-text-ic) compares the scores on the test half; [Figure 13.1](#fig-ml-text-learning) shows how each depends on the number of labelled headlines.

| score | rank IC with the reaction return |
| --- | --- |
| dictionary (11 positive, 16 negative words) | 0.224 |
| tf-idf words, [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) | 0.273 |
| tf-idf words and bigrams, [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) | 0.257 |
| return-supervised word list | 0.238 |
| [word embeddings](#def-ml-text-from-bag-of-words-to-embeddings-embedding) (co-occurrence SVD), ridge regression | 0.094 |
| event rules, mean training return per event | 0.292 |
| the planted effect | 0.295 |

***Table 13.1.** Rank information coefficients on the 8 077 test headlines, models trained on the 11 923 earlier ones. Data: `ml_text.scores`.*

![Rank IC on the test half against the number of labelled headlines each model was trained on (the dictionary uses none). With 300 headlines no word passes the supervised screen and its score is zero. Data: ml_text.learning_curve.](https://one-course.com/images/onecourse/chapters/quant-12/ml-text-from-bag-of-words-to-embeddings/fig-76c1f1917f52.svg)

***Figure 13.1.** Rank IC on the test half against the number of labelled headlines each model was trained on (the dictionary uses none). With 300 headlines no word passes the supervised screen and its score is zero. Data: `ml_text.learning_curve`.*

Three lessons come out of the curves. A dictionary costs no labels: the tf-idf word model passes it only between 1 000 and 3 000 labelled headlines, the supervised word list and the bigram model only beyond 3 000, and for real text, with a far larger vocabulary and weaker signals, enough labels means years of news. The [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) on tf-idf words (chapter 4’s model, with a ridge penalty) ends highest of the bag-of-words scores, at 0.273 against the dictionary’s 0.224. And bigrams, which can represent the compound headlines a word model cannot, cost more in variance than they gained: 0.257 against 0.273, with nearly four times the features and the compounds only 9% of the headlines. A richer representation needs more data to pay for itself.

## 13.3 Topics and events

**Definition 13.2 (Topic model, latent Dirichlet allocation).**

A *topic model* describes each document as a mixture of a few topics, each topic a distribution over words, all estimated from word counts without labels. *Latent Dirichlet allocation* (LDA) is the Bayesian topic model in which each document’s topic shares and each topic’s word probabilities have Dirichlet priors, and each word is drawn by first drawing its topic from the document’s shares (Blei, Ng and Jordan, 2002).

| topic | six most probable words |
| --- | --- |
| 1 | analysts, fell, short, expectations, forecasts, exceeded |
| 2 | outlook, guidance, but, expectations, forecasts, estimates |
| 3 | annual, meeting, record, attendance, reports, holds |
| 4 | full, forecast, year, reports, line, costs |
| 5 | wins, authorises, program, repurchase, industry, award |
| 6 | announces, gets, rating, earnings, schedules, call |

***Table 13.2.** Six topics fitted by LDA to the training headlines (the company placeholder, prepositions and context words removed). Data: `ml_text.topics`.*

[Table 13.2](#tab-ml-text-topics) shows what a [topic model](#def-ml-text-from-bag-of-words-to-embeddings-lda) finds in six-word headlines: the corpus’s templates, not its news. Beats and misses share topic 1, guidance changes and compounds of either direction share topic 2, and a buyback shares topic 5 with an industry award; on average only 63% of an event’s headlines fall in its most common topic. A topic is a pattern of co-occurrence, and words of opposite meaning co-occur with the same words. Topics are useful to organise long documents (which sections of a filing changed, what an earnings call spent its time on) and as features for a supervised model, not as a sentiment.

**Definition 13.3 (Event extraction).**

*Event extraction* turns text into structured records (who, what kind of event, which direction, when), by rules on keywords and syntax or by a trained tagger, so that a model can be fitted on events rather than on words.

The chapter’s extractor is a dozen rules: a keyword per event, “not” before a guidance cut, a result and a guidance change joined by “but”. It reads every one of the 20 000 headlines correctly, and a score that assigns each extracted event its mean training return reaches 0.292 against the planted effect’s 0.295, above every bag-of-words model. It should: its rules were written from the corpus’s own templates. On a headline it was not written for (“surpasses consensus”) it returns nothing. Rules are precise and brittle; learned models are robust and hungry for labels; production systems use rules for the events that matter most and learned taggers for the rest, and measure both on text they were not built from.

## 13.4 Embeddings

**Definition 13.4 (Word embedding).**

A *word embedding* maps each word to a dense vector of a few dozen to a few thousand real numbers, learned so that words used in similar contexts get nearby vectors: by factorising a matrix of co-occurrence statistics, or by training a network to predict a word from its neighbours (Mikolov and co-authors, 2013).

The chapter’s embeddings are the first kind ([Listing 13.2](#lst-ml-text-ppmi)): the positive pointwise mutual information of words appearing in the same headline, reduced to 16 dimensions by a singular value decomposition (Book 4, chapter 25), rows scaled to unit length. They learn exactly what the definition promises. Synonyms that fill the same slot, “beat” and “tops”, “raised” and “lifts”, “guidance” and “outlook”, have cosine similarities above 0.99; so do antonyms that fill the same slot, “beat” and “missed” at 0.95 ([Figure 13.2](#fig-ml-text-pairs)). Context decides similarity, and a word and its opposite share their contexts. A ridge regression on the average of a headline’s word vectors then has little to work with: IC 0.094. Embeddings trained on a large general corpus carry more than this toy’s, but the property is the same, and it is why sentiment models built on embeddings are fine-tuned on labels (chapter 10) and why chapter 14’s language models are evaluated on the task, not on their vectors.

![Cosine similarities of word embeddings learned from the training headlines’ co-occurrences. Opposites that fill the same slot are as close as synonyms. Data: ml_text.pairs.](https://one-course.com/images/onecourse/chapters/quant-12/ml-text-from-bag-of-words-to-embeddings/fig-caa9591a2701.svg)

***Figure 13.2.** Cosine similarities of [word embeddings](#def-ml-text-from-bag-of-words-to-embeddings-embedding) learned from the training headlines’ co-occurrences. Opposites that fill the same slot are as close as synonyms. Data: `ml_text.pairs`.*

## 13.5 Entity linking to securities

**Definition 13.5 (Named-entity recognition, entity linking).**

*Named-entity recognition* finds the spans of a text that name entities (companies, people, places, products) and their types. *Entity linking* maps each recognised span to one identifier in a knowledge base; for a trading firm, to a permanent identifier in the security master (Book 7, chapter 4) as of the document’s date.

A model that reads news perfectly and trades the wrong company is worse than no model. The chapter’s recogniser is a gazetteer: every name, short name and ticker any company has held, matched longest first, which finds every mention in the corpus because the corpus uses no other names; real recognisers also tag names never seen before, from capitalisation, syntax and a trained tagger. The linker ([Listing 13.3](#lst-ml-text-link)) then resolves a ticker through the security master and a name through an alias table, both as of the headline’s date, and separates several candidates by sector words in the headline, then by coverage.

|  | aliases as of the headline’s date | today’s aliases |
| --- | --- | --- |
| mention | correct | linked | precision | correct | linked | precision |
| full name | 100.0 | 100.0 | 100.0 | 89.3 | 89.3 | 100.0 |
| short name | 98.2 | 100.0 | 98.2 | 86.8 | 88.6 | 98.0 |
| ticker | 100.0 | 100.0 | 100.0 | 84.5 | 92.4 | 91.5 |
| all | 99.5 | 100.0 | 99.5 | 87.5 | 89.7 | 97.5 |

***Table 13.3.** [Entity linking](#def-ml-text-from-bag-of-words-to-embeddings-ner) on the 20 000 headlines (%): share linked to the right company, share linked to any company, and precision (right among linked). Data: `ml_text.linking`.*

[Table 13.3](#tab-ml-text-linking) holds the chapter’s second result. With point-in-time aliases the linker is right 99.5% of the time; its only errors are the short name Apex in the 245 headlines that carry no sector word, where it falls back on the most covered Apex and is right 58.0% of the time (with a sector word it is right in all 737). With today’s aliases it is right 87.5% of the time: old names and old tickers resolve to nothing, and the four reused tickers resolve to the new listing, so that 8.5% of ticker mentions it links go to the wrong company, silently. A backtest on those links would trade a takeover headline about one company in the shares of another; the rule of Book 7, chapter 3, applies to aliases as to prices.

**Method 13.6 (A text pipeline for trading).**

1. Recognise mentions and link them point in time to permanent identifiers; measure linking precision on a labelled sample, by mention kind.
2. Mask the mentions, tokenise, and build the representation (tf-idf words first).
3. Start with a dictionary as the baseline; fit supervised scores on returns in walk-forward folds, with the reaction window and horizon fixed in advance.
4. Extract the events that matter with rules or a tagger and test them against the bag-of-words model.
5. Treat embeddings and topics as features to be validated, not as sentiment.

## 13.6 Tutorial: which Apex?

**Goal.** Generate the corpus, recognise and link mentions point in time, and compare five text scores by their information coefficient. **End state:** Tables [13.1](#tab-ml-text-ic), [13.2](#tab-ml-text-topics) and [13.3](#tab-ml-text-linking), Figures [13.1](#fig-ml-text-learning) and [13.2](#fig-ml-text-pairs).

1. **A return-supervised word list.** `def supervised_words (docs, returns, min_count=30 , z=3.0 ): """Keep the words whose documents' mean return differs from the rest's by more than z standard errors; each word's weight is that difference (the return-supervised screen, in the spirit of Ke, Kelly and Xiu).""" r = np.asarray(returns, dtype=float ) where = defaultdict(list ) for i, toks in enumerate (docs): for w in set (toks): where[w].append(i) sd, n, tot = r.std(), len (r), r.sum() out = {} for w, idx in where.items(): k = len (idx) if k < min_count or k > n - min_count: continue m_in = r[idx].mean() m_out = (tot - r[idx].sum()) / (n - k) se = sd * np.sqrt(1 / k + 1 / (n - k)) if abs (m_in - m_out) > z * se: out[w] = float (m_in - m_out) return out` **Listing 13.1.** A return-supervised word screen. code/firm/textml/firm_textml.py
2. **Embeddings from co-occurrence.** `def ppmi_svd (docs, k=16 , min_count=20 ): """Word vectors: positive pointwise mutual information of words co-occurring in a document, reduced by a truncated singular value decomposition; rows scaled to unit length.""" cnt = Counter(t for d in docs for t in set (d)) vocab = sorted (w for w, c in cnt.items() if c >= min_count) ix = {w: i for i, w in enumerate (vocab)} C = np.zeros((len (vocab), len (vocab))) for d in docs: ids = sorted ({ix[t] for t in d if t in ix}) for a in ids: for b in ids: if a != b: C[a, b] += 1.0 tot = C.sum() row = C.sum(1 , keepdims=True ) with np.errstate(divide=" ignore " , invalid=" ignore " ): pmi = np.log(C * tot / (row @ row.T)) M = np.where(np.isfinite(pmi) & (pmi > 0 ), pmi, 0.0 ) U, S, _ = np.linalg.svd(M) V = U[:, :k] * np.sqrt(S[:k]) V /= np.linalg.norm(V, axis=1 , keepdims=True ) + 1e-12 return vocab, V` **Listing 13.2.** Word embeddings from positive pointwise mutual information and an SVD. code/firm/textml/firm_textml.py
3. **Point-in-time linking.** `def link (surface, date, context, table, sm, sectors, prior): """Map a recognised surface to a permanent id as of `date`: a ticker through the security master, a name or a short name through the alias table; several candidates are separated by sector words in the context, then by the prior (the candidate with most past coverage).""" if surface.isupper(): return sm.resolve(surface, date) cands = sorted ({pid for _, pid in table.valid(surface, date)}) if len (cands) > 1 : ctx = set (context) by_sector = [p for p in cands if ctx & set (sectors[p])] cands = by_sector if len (by_sector) == 1 else sorted (cands, key=lambda p: -prior.get(p, 0 )) return cands[0 ] if cands else None` **Listing 13.3.** Linking a mention to a permanent identifier as of a date. code/firm/textml/firm_textml.py
4. **Run** `ml_text.scores()` , `learning_curve()` , `topics()` , `pairs()` , `linking()` and `fig_text.py` .

**What to change next.** Add phrasings to the test period that the training period never saw, and watch the rules and the supervised models fail differently; give the linker a context model instead of a sector-word list.

## 13.7 Build: text features and entity linking

**Purpose.** Text turned into scores that can be validated like any other feature, attached to the right security.

**Interface.** `tokenize`, `ngrams`, `dictionary_score(tokens, pos, neg)`, `supervised_words(docs, returns, min_count, z)`, `supervised_score`, `ppmi_svd(docs, k, min_count)`, `doc_vectors`, `extract_event(tokens)`; `AliasTable`, `recognise(text, gazetteer)`, `link(surface, date, context, table, sm, sectors, prior)`; `news_corpus(n_docs, seed, days)` on `firm.secmaster`.

**Rules.** Mentions are masked before any model sees the text; aliases and tickers are resolved as of the document’s date; supervised scores are fitted on earlier documents than they score.

**Acceptance tests.** `code/firm/textml/tests/`: tokens and [n-grams](#def-ml-text-from-bag-of-words-to-embeddings-bow) of a known headline; the supervised screen finds a planted word and ignores a neutral one; embeddings put words with identical contexts together; the rules read compounds and negations; point-in-time linking follows a ticker change, a rename and a reused ticker, and today’s table does not.

**Stretch.** A trained sequence tagger for recognition; linking by a context model over the candidates’ descriptions; the Ke–Kelly–Xiu [topic model](#def-ml-text-from-bag-of-words-to-embeddings-lda) on top of the screen.

Sources and further reading

- P. C. Tetlock, “Giving content to investor sentiment: the role of media in the stock market”, *Journal of Finance* 62(3), 2007.
- T. Loughran and B. McDonald, “When is a liability not a liability? Textual analysis, dictionaries, and 10-Ks”, *Journal of Finance* 66(1), 2011.
- Z. T. Ke, B. T. Kelly and D. Xiu, “Predicting returns with text data”, NBER Working Paper 26186, 2019.
- M. Gentzkow, B. Kelly and M. Taddy, “Text as data”, *Journal of Economic Literature* 57(3), 2019.
- D. M. Blei, A. Y. Ng and M. I. Jordan, “Latent Dirichlet allocation”, *Advances in Neural Information Processing Systems* 14, MIT Press, 2002.
- T. Mikolov, K. Chen, G. Corrado and J. Dean, “Efficient estimation of word representations in vector space”, arXiv:1301.3781, 2013.

## 13.8 Exercises

**Exercise 13.1 ★.**

List the tokens and the bigrams of “`<co>` did not cut guidance”. Which bigram carries the negation?

**Solution of Exercise 13.1.**

Tokens: `<co>`, did, not, cut, guidance. Bigrams: `<co> did`, did not, not cut, cut guidance. “Not cut” carries the negation; the word “cut” alone reads as bad news.

**Exercise 13.2 ★.**

In a corpus of 20 000 documents, “guidance” appears in 3 000 and “halves” in 200. What are their inverse document frequencies ($\log N/n_w$)? Which weighs more in a document containing both once?

**Solution of Exercise 13.2.**

$\log(20\,000/3\,000) = 1.90$ and $\log(20\,000/200) = 4.61$: “halves” weighs about 2.4 times as much. Rare words are more informative about which document this is, which is not the same as more informative about the return.

**Exercise 13.3 ★.**

Compute the dictionary score of “`<co>` beat estimates but cut guidance” (six tokens) and of “`<co>` wins industry award” (four tokens) with the chapter’s lists. What are the planted effects?

**Solution of Exercise 13.3.**

$(1 - 1)/6 = 0$ for the compound (“beat” positive, “cut” negative), planted $-1.2\%$; $1/4 = 0.25$ for the award (“wins” positive), planted 0.

**Exercise 13.4 ★★.**

Why do “beat” and “missed” get almost the same embedding here, and what would separate them?

**Solution of Exercise 13.4.**

They appear in the same slot with the same neighbours (a company, then “estimates”, “forecasts” or “expectations”), and co-occurrence is all the embedding sees: cosine 0.95. Only a signal that differs between them separates them: returns (a supervised fine-tune), or contexts that differ (“beat … shares rise”), which a larger corpus supplies in part.

**Exercise 13.5 ★★.**

The supervised screen keeps “to” and “be”. Is that a bug? When would it become one?

**Solution of Exercise 13.5.**

Not a bug here: in this corpus “to be” appears only in takeover headlines, so it carries their return. It becomes one when the phrasing that made it informative changes (a new wire’s style, a new template), or when the word stands for something outside the text, such as one outlet’s coverage of one kind of company; a screened word should be read, and the list checked for stability across folds.

**Exercise 13.6 ★★.**

*Find the flaw.* “We linked five years of news to tickers with the vendor’s current symbol file; 97% of the links we checked by hand were right, so linking is not a problem.”

**Solution of Exercise 13.6.**

Today’s symbol file maps old tickers to nothing and reused tickers to today’s holder; the errors concentrate in the past, in companies that changed or disappeared, which is exactly where survivorship and takeover effects live. A hand check of recent links, or of links that resolved, measures the wrong population. Here today’s aliases give 97.5% precision among linked mentions but only 87.5% correct overall, and 91.5% precision on tickers. Check a sample stratified by year and by mention kind, including the unlinked ones, against point-in-time aliases.

**Exercise 13.7 ★★★.**

*Coding.* Train the supervised screen and the tf-idf [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) on the training half with the company mentions left in the text (unmasked). What do the models learn about the companies, and what happens to the test IC? Explain.

**Solution of Exercise 13.7.**

Unmasked, the tf-idf model’s IC falls from 0.273 to 0.266: the company words get small weights fitted to noise, since no company effect was planted. The supervised screen keeps no company word and scores 0.238 either way, because its three-standard-error bar rejects them. On real data company words would carry each company’s returns over the training period (a good year, a takeover), a leak of past performance dressed as text, which masking prevents.

**Exercise 13.8 ★★★.**

Show that with a document’s tokens as a [bag of words](#def-ml-text-from-bag-of-words-to-embeddings-bow), no linear model on word counts can give “beat … but cut” an effect of $-1.2$ and “missed … but raised” an effect of $+1.2$ while giving “beat”, “missed”, “raised” and “cut” alone $+1$, $-1$, $+1.5$ and $-1.5$. What features would?

**Solution of Exercise 13.8.**

For a linear model with intercept $b$ and word weights $w$, the compound “$A$ but $C$” scores $b + w_A + w_{\text{but}} + w_C
= f(A) + f(C) + (w_{\text{but}} - b)$, where $f(A)$ and $f(C)$ are the scores of the single-clause headlines with the same words. Matching $f(\text{beat}) = 1$, $f(\text{cut}) = -1.5$ and $f(\text{beat but cut}) = -1.2$ requires $w_{\text{but}} - b =
-0.7$; matching $f(\text{missed}) = -1$, $f(\text{raised}) = 1.5$ and $f(\text{missed but raised}) = 1.2$ requires $+0.7$. The two conditions contradict each other. Bigrams such as “but cut” and “but raised”, or an interaction between the clause after “but” and a guidance indicator, can fit both: the second clause dominates.

## 13.9 Problem: Which Apex?

**Problem 13.1.**

Weekend problem — text to the right trade

The chapter’s corpus: 20 000 headlines, 32 companies (three called Apex), six ticker changes, three renames, four reused tickers.

**Part I — Representation.**

1. Why mask the company mentions before building features?
2. What does tf-idf do to a word that appears in every headline?
3. How many features do bigrams add, roughly, and why did they not help here?
4. What does the planted effect’s IC of 0.295 say about the best any text model can do?

**Part II — Scores.**

5. Rank the six scores of [Table 13.1](#tab-ml-text-ic) .
6. Below how many labelled headlines is the dictionary the best score?
7. Why do the event rules beat every learned model, and why should you not believe it?
8. Why do embeddings score 0.094?

**Part III — Linking.**

9. What share of mentions does point-in-time linking get right, and where are its errors?
10. What does today’s table get right, and what are its two failure modes?
11. Which failure is worse for a backtest, an unlinked mention or a wrongly linked one?
12. How accurate is the linker on Apex with and without a sector word?

**Part IV — The verdict.**

13. State the *named result* : entity-linking accuracy and precision with and without point-in-time aliases, and the IC of the return-supervised models against the dictionary’s.
14. What would you build first for a news desk: better sentiment or better linking?
15. How would you check a vendor’s entity links?
16. How do you keep a supervised text model honest in time?
17. What is a [topic model](#def-ml-text-from-bag-of-words-to-embeddings-lda) good for in this business?
18. What changes with long documents (filings, transcripts) rather than headlines?
19. What would make the dictionary competitive again?
20. In one sentence: what is the most important step between a headline and a trade?

**Solution of Problem 13.1.**

**Part I.**

1. So that no model learns a company’s identity (its past returns in the sample) as a text signal, and so that one model serves every company.
2. Gives it weight $\log(N/N) = 0$ : it cannot distinguish documents.
3. From 123 words to 473 features, nearly four times as many; the compounds they could capture are 9% of the headlines, and the added variance cost more than the fit gained.
4. The planted effect explains all that text can: an IC of about 0.3, because the reaction’s noise (3%) is large against the effects (up to 2%). Any score above that on a [test set](https://one-course.com/books/quant/12/en/chapter/1-why-financial-machine-learning-is-different#def-ml-why-financial-machine-learning-is-different-sets) is luck or leakage.

**Part II.**

1. Event rules 0.292, tf-idf words 0.273, tf-idf with bigrams 0.257, supervised words 0.238, dictionary 0.224, embeddings 0.094.
2. The tf-idf word model passes the dictionary between 1 000 and 3 000 labelled headlines; the supervised list and the bigram model only beyond 3 000.
3. They were written from the corpus’s own templates and read all 20 000 headlines correctly; they fail on a phrasing they were not written for (“surpasses consensus” returns nothing). Measure them on text written after the rules.
4. The embeddings learned from co-occurrence put opposites together (“beat” and “missed” at 0.95), so the average of a headline’s vectors barely distinguishes good from bad news.

**Part III.**

1. 99.5%; its only errors are the short name Apex with no sector word, where the most covered Apex is chosen.
2. 87.5%. Old names and tickers resolve to nothing (89.7% of mentions linked), and reused tickers resolve to the new holder (precision 91.5% on tickers).
3. The wrong link: an unlinked headline is a missed trade, a wrongly linked one is a trade in the wrong security, with a return that looks like noise or, worse, like a signal.
4. 100% of 737 with a sector word, 58.0% of 245 without.

**Part IV.**

1. *Which Apex?* Point-in-time linking is right on 99.5% of the mentions (precision 99.5%); linking with today’s aliases is right on 87.5% (precision 97.5% overall, 91.5% on tickers). The return-supervised models reach an IC of 0.273 (tf-idf words, [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) ) and 0.238 (the supervised word list) against the dictionary’s 0.224.
2. Linking: every model downstream inherits its errors, and its errors are silent.
3. Sample links stratified by year and mention kind, including the unlinked, and compare with point-in-time aliases; count links to delisted or renamed companies.
4. Fit and select it in walk-forward folds on earlier text, fix the reaction window in advance, and re-screen on a schedule with the list’s stability checked.
5. Organising long documents and building features for a supervised model; not measuring tone.
6. Tokens per document rise to thousands, dictionaries become more stable, section structure and changes from the previous filing carry information, and a mention may concern a peer, a supplier or a customer rather than the filer.
7. Few labels, a new domain, or a need for a transparent score; and lists built for the domain.
8. Getting the security right.

## 13.10 Interview questions

**Interview question 13.1 ★ researcher, mle.**

What is tf-idf, and why does it help a linear model on text?

**Solution of Interview question 13.1.**

Term counts weighted by the log inverse share of documents containing the term, rows normalised. It shrinks common words towards zero and puts documents of different lengths on one scale, so a penalised linear model spends its budget on distinctive words.

*What the interviewer is looking for: the formula, the reason for the idf, and length normalisation.*

**Interview question 13.2 ★ researcher.**

Why do finance-specific dictionaries exist?

**Solution of Interview question 13.2.**

General lists call common financial words negative (liability, tax, cost, capital) and miss financial meanings; Loughran and McDonald (2011) found that almost three-quarters of the negative words of a widely used general dictionary are not negative in 10-Ks, and built finance lists.

*What the interviewer is looking for: misclassification of domain words, with the reference or an example.*

**Interview question 13.3 ★★ researcher.**

How would you build a sentiment score supervised by returns, and what could leak into it?

**Solution of Interview question 13.3.**

Screen words by their association with subsequent returns on training data, weigh them, score documents; or fit a penalised [logistic regression](https://one-course.com/books/quant/12/en/chapter/4-linear-and-regularised-baselines#def-ml-linear-and-regularised-baselines-logistic) on tf-idf. Leaks: returns before the text’s timestamp (use the time the text became available), company names standing for their returns, labels from overlapping windows, and vocabulary chosen on the full sample.

*What the interviewer is looking for: a supervised construction and at least two concrete leaks with remedies.*

**Interview question 13.4 ★★ mle.**

Why are antonyms close in word-embedding space, and what does it mean for using embeddings as features?

**Solution of Interview question 13.4.**

Embeddings are trained to reflect shared contexts, and a word and its opposite fill the same contexts. So distances in the space measure topic and syntax, not polarity; as features they need a supervised layer or [fine-tuning](https://one-course.com/books/quant/12/en/chapter/10-representation-learning#def-ml-representation-learning-ssl) on the target.

*What the interviewer is looking for: the distributional reason and the consequence for sentiment.*

**Interview question 13.5 ★★ mle, developer.**

Design the entity-linking step for a news feed that drives trading. What can go wrong?

**Solution of Interview question 13.5.**

Recognise mentions (gazetteer and a tagger), resolve candidates from point-in-time aliases and tickers, disambiguate by context, output a permanent identifier with a confidence and abstain below a threshold; log everything. Failure modes: ticker reuse, renames, shared short names, subsidiaries and brands, and a symbol file refreshed without history.

*What the interviewer is looking for: point-in-time resolution, disambiguation, abstention, and named failure modes.*

**Interview question 13.6 ★★★ researcher.**

Your text model’s IC fell by half in its first live year. List the causes you would check, in order.

**Solution of Interview question 13.6.**

First the plumbing: timestamps (text time against trade time), linking changes, a vendor’s format or coverage change. Then leakage in the backtest: vocabulary or labels chosen on the full sample, the reaction window. Then drift: new phrasings, crowding of the signal (Book 7, chapter 28), and finally the expected shrinkage of an overfitted model.

*What the interviewer is looking for: an ordering from data problems to leakage to genuine decay.*
