Machine Learning for Markets · Machine learning
14Large Language Models in Finance
A language model asked whether headlines from its own training period were good news for the stock is right 70% of the time; asked about headlines written after its training data end, it is right 56% of the time. The first number measures its memory, the second its judgement, and a backtest that mixes the two measures nothing. The model in question is this chapter’s: a transformer of 107 000 parameters trained in a minute on the synthetic headlines of chapter 13, small enough to be retrained with and without the evaluation period in its data. The protocol it tests is the one a desk needs for commercial models whose training data it cannot see. The chapter then turns to what language models are used for in practice, reading filings and news at scale through retrieval, and running research tasks as agents, and to what a desk can and cannot hand them. Nothing in it calls a network service or downloads a model; every number comes from local code with tests.
14.1 Language models and how they are trained
Definition 14.1 (Language model, large language model)
A language model assigns probabilities to sequences of tokens, usually through the conditional probability of each token given the ones before it, ; it generates text by sampling one token at a time. A large language model (LLM) is a transformer language model (chapter 8) with billions of parameters, pre-trained (chapter 10) on a large share of the public text of the internet and books, then fine-tuned to follow instructions.
Definition 14.2 (Prompt, context window, training cut-off)
A prompt is the text given to a language model, which it continues. The context window is the largest number of tokens the model can condition on at once, prompt and answer together. The training cut-off is the date of the latest text in the model’s training data: the model may know anything published before it and knows nothing published after it, except what the prompt tells it.
Listing 14.1 is the architecture at toy scale: token and position embeddings, two layers of causal self-attention (a token attends only to earlier tokens), and a projection back onto the 505 words of the vocabulary. It is trained, like its large relatives, to predict each next token. Its training text is the chapter-13 corpus, cut to its first 10 000 headlines and written as sentences such as w190 apex mining beat estimates amid copper demand => up: a week, the company, the headline, and after the arrow what the stock did (up or down, the sign of the reaction return). Asking the model about a headline is prompting it with everything up to the arrow and comparing the probabilities it gives to up and down. Two copies are trained: a clean model whose training data end on day 907 and a contaminated model whose data run to day 1 210; both are scored on the 5 972 headlines before day 907, the 2 016 between the two cut-offs and the 2 012 after both.
Large models are built the same way at a scale this chapter cannot reach: BloombergGPT (Wu and co-authors, 2023), for instance, is a 50-billion-parameter model trained on a 363-billion-token financial dataset built from Bloomberg’s sources together with 345 billion tokens of general text. What carries over from the toy is the protocol, not the numbers.
14.2 Filings, calls and news at scale: retrieval
A language model knows only its training data and its prompt. To answer questions about a document collection that changes every day (filings, transcripts, news), it is given the relevant passages in its prompt.
Definition 14.3 (Retrieval-augmented generation, BM25)
Retrieval-augmented generation (RAG) answers a question by first retrieving passages relevant to it from a document collection and then generating the answer from a prompt that contains them (Lewis and co-authors, 2020). BM25 is the standard lexical retrieval score: for each query term found in a passage, the term’s inverse document frequency times a saturating function of its count in the passage, normalised by the passage’s length (Robertson and Zaragoza, 2009).
The chapter’s collection holds 540 passages from 135 synthetic annual filings, one passage per company, fiscal year and metric (revenue, net debt, headcount, capital expenditure), each written with one of three synonyms (“revenue”, “sales” or “turnover”) and two context words. One question is asked about each passage, worded with a synonym drawn at random, so that about a third use the passage’s own word; the answer is extracted from the top passage, so that a right passage gives the right number. Figure 14.1 compares four searches. BM25 (Listing 14.2) finds every passage asked about in its own words and 5.0% of the others: a question about “sales” goes to another company’s passage that says “sales”. Dense retrieval, the cosine between the average word embeddings of question and passage, knows that “sales” means revenue (the embeddings, trained on 5 000 general business sentences, put the three synonyms at cosine above 0.99) but not which company is asked about: 1.1%. Adding the two standardised scores helps little (18.2%). Expanding the query with its words’ embedding neighbours before running BM25 finds 89.1% of the paraphrased questions and 86.3% of the others: 88.1% overall.
ml_llm.retrieval.Definition 14.4 (Hallucination)
A hallucination is a fluent answer that is not supported by the model’s input or by fact: a figure, a quotation or a reference produced because it is probable text, not because it is true.
A retrieval system that always answers hallucinates by construction on questions its collection cannot answer. On 300 questions about company-years with no filing, the expanded search returns a passage and a number every time. A grounding check, answering only when the passage names the question’s company and fiscal year, abstains on all 300; on the answerable questions it turns away 11.9%, exactly the 64 of 540 whose top passage was about another company or year, which would have been wrong. A generative model on top of retrieval needs the same discipline: quote the passage, check the entities, and abstain rather than guess.
14.3 Evaluation without leakage from the training cut-off
Definition 14.5 (Memorisation, anonymisation test)
Memorisation is a model’s ability to reproduce specific training examples, including their outcomes, rather than a rule that generalises; language models memorise text seen even once (Carlini and co-authors, 2021). The anonymisation test replaces the identifiers in an input (company names, tickers, dates) by placeholders and measures how much the model’s accuracy falls: a model that reads the text loses little, a model that recalls the event loses its memory.
Evaluating a commercial language model on historical headlines is look-ahead bias (Book 7, chapter 3) in a new form: the model may have read the headline and what the stock did next. Figure 14.2 measures both tests on the chapter’s two models. On the headlines before day 907, which both saw with their outcomes, the clean model is right 69.7% of the time and the contaminated one 65.4% with names and weeks, and 57.7% and 57.2% without: that gap is memory. Between the two cut-offs the contaminated model is right 64.8% of the time and the clean model 56.2%, and anonymisation takes the contaminated model to 57.0% and leaves the clean one where it was (56.0%). After both cut-offs the contaminated model is right 59.2% of the time, the clean one 56.3%, both below the 61.6% that the sign of the planted effect achieves.
ml_llm.contamination.The chapter’s named result follows. The memorisation premium, the accuracy between the cut-offs minus the accuracy after both, is 5.6 points for the contaminated model and for the clean one; the anonymisation loss between the cut-offs is 7.8 points for the contaminated model and 0.2 for the clean one. Both tests catch the contamination without knowing the training data, which is their purpose: a desk evaluating a model whose cut-off it doubts runs them, and trusts only numbers from after the stated cut-off. The literature found the same two effects in real models. Lopez-Lira and Tang (2023) evaluated GPT-4 on headlines published after its knowledge cut-off for that reason; Glasserman and Lin (2023) compared original headlines with headlines stripped of company identifiers and found, in-sample, that the anonymised ones did better: general knowledge of the company distracted the model more than look-ahead helped it. The toy has no general knowledge to be distracted by, and after both cut-offs anonymisation costs each model about a point (59.2% to 58.1%, 56.3% to 55.3%): the anonymisation test is a test of memory, and its loss must be read against that small cost.
Method 14.6 (Evaluating a language model on market data)
- Establish the training cut-off from the provider’s documentation and treat everything before it as in-sample.
- Score only on text published after the cut-off; if history before it must be used, report the anonymisation test and the accuracy by period, and expect them to disagree.
- Freeze the model version, the prompt and the decoding settings in the research log (Book 7, chapter 1): a provider’s silent update is a new model.
- Compare against a cheap baseline (a dictionary, a tf-idf model) on the same headlines and horizon.
14.4 Agents in research workflows
Definition 14.7 (Language-model agent, tool call)
A language-model agent is a loop in which a language model chooses actions, observes their results and decides the next action until a task is done (Yao and co-authors, 2023). A tool call is one such action: the model emits a structured request (a function name and arguments) that the surrounding program executes, returning the result to the model.
An agent that can search, compute and run backtests is a research assistant with the same failure modes as a junior analyst, at machine speed: it will use data it should not have. Listing 14.3 is the firm’s harness: tools are registered by name, a dated tool receives the as-of date of the task and any result published after it is refused, and every call, allowed or refused, goes to the log. The chapter’s agent is a fixed plan rather than a language model, so that its behaviour can be tested: asked, as of a random day, for the latest annual figure of a company, it searches, keeps the passages about the company and the metric, and answers from the latest fiscal year it finds. On 400 such questions, without the guard, 80.8% of its answers come from filings published after the as-of date and 19.2% are right; with the guard, none do and all are right, and each question leaves one logged call.
14.5 What a desk can and cannot delegate
Language models are good at the work between data and decision that is expensive for people and cheap to check: extracting fields from filings, summarising a call against the previous one, drafting code and tests, tagging news with events and entities, answering questions over a collection with the passage quoted. Each of those has a checkable output. What cannot be delegated is what cannot be checked cheaply: the decision to trade, the judgement that a backtest is clean, the claim that a number is right when no passage supports it. Three rules keep the boundary. Nothing the model returns enters a position or a report without a source that a person or a test can verify. Material non-public information (Book 7, chapter 12) stays out of prompts sent to services the firm does not control, since a prompt is a disclosure. And every model call that feeds research is logged with its version, prompt and output, so that the result can be reproduced or at least explained. Kim, Muhn and Nikolaev (2024) report that GPT-4, given standardised and anonymous financial statements, predicted the direction of earnings changes better than analysts; the anonymisation is what made the result credible.
14.6 Tutorial: the model that knew
Goal. Train a clean and a contaminated language model on the chapter-13 headlines and measure the memorisation premium and the anonymisation loss; compare four retrieval methods and a grounding check; run the agent with and without its guard. End state: Figures 14.2 and 14.1.
A decoder-only language model at toy scale.
class TinyLM(nn.Module): """Token and position embeddings, a stack of causal self-attention layers, and a projection back onto the vocabulary: the architecture of a decoder-only language model at toy scale.""" def __init__(self, vocab_size, d=64, heads=4, layers=2, max_len=32): super().__init__() self.tok = nn.Embedding(vocab_size, d) self.pos = nn.Embedding(max_len, d) layer = nn.TransformerEncoderLayer(d, heads, dim_feedforward=4 * d, dropout=0.0, batch_first=True) self.blocks = nn.TransformerEncoder(layer, layers, enable_nested_tensor=False) self.out = nn.Linear(d, vocab_size) def forward(self, ids): # ids: (m, T) -> logits (m, T, vocab) T = ids.shape[1] mask = torch.triu(torch.full((T, T), float("-inf")), diagonal=1) # a token sees only its past z = self.tok(ids) + self.pos(torch.arange(T))[None] return self.out(self.blocks(z, mask=mask, is_causal=True))Listing 14.1. A causal transformer language model. code/firm/llmeval/firm_llmeval.py BM25.
class BM25: """score(q, d) = sum over query terms t of idf(t) * f(t,d) (k1 + 1) / (f(t,d) + k1 (1 - b + b |d| / avgdl)), idf(t) = log(1 + (N - n_t + 0.5) / (n_t + 0.5)).""" def __init__(self, docs, k1=1.5, b=0.75): self.docs = [Counter(d) for d in docs] self.len = np.array([len(d) for d in docs], dtype=float) self.avg = self.len.mean() self.k1, self.b = k1, b df = Counter(t for d in docs for t in set(d)) N = len(docs) self.idf = {t: math.log(1 + (N - n + 0.5) / (n + 0.5)) for t, n in df.items()} def scores(self, query): s = np.zeros(len(self.docs)) norm = self.k1 * (1 - self.b + self.b * self.len / self.avg) for t in set(query): if t not in self.idf: continue f = np.array([d.get(t, 0) for d in self.docs], dtype=float) s += self.idf[t] * f * (self.k1 + 1) / (f + norm) return sListing 14.2. The BM25 score. code/firm/llmeval/firm_llmeval.py Point-in-time tools.
class ToolRegistry: """Tools called by name; a dated tool receives the registry's as-of date and must not return anything published after it. Every call, its arguments, its result size and any refusal go to the log.""" def __init__(self, as_of): self.as_of, self.tools, self.log = as_of, {}, [] def register(self, name, fn, dated=False): self.tools[name] = (fn, dated) def call(self, name, **kw): if name not in self.tools: self.log.append({"tool": name, "args": kw, "status": "unknown tool"}) return None fn, dated = self.tools[name] if dated and kw.get("as_of", self.as_of) > self.as_of: self.log.append({"tool": name, "args": kw, "status": "refused: after the as-of date"}) return None args = {"as_of": self.as_of} | kw if dated else kw out = fn(**args) if dated and any(r.get("published", -math.inf) > self.as_of for r in (out or [])): self.log.append({"tool": name, "args": kw, "status": "refused: result after the as-of date"}) return None self.log.append({"tool": name, "args": kw, "status": "ok", "n": len(out) if hasattr(out, "__len__") else 1}) return outListing 14.3. A tool registry with a point-in-time guard and a log. code/firm/llmeval/firm_llmeval.py - Run
ml_llm.contamination()(about a minute on one core),verdict(),retrieval(),unanswerable(),agent_run()andfig_llm.py.
What to change next. Train the contaminated model for fewer epochs and watch the premium shrink; swap company names between headlines (an entity-swap test) instead of removing them.
14.7 Build: evaluating language models without leakage
Purpose. Language models and retrieval used where their output can be checked, and evaluated only where their training data cannot have seen the answer.
Interface. Vocab, TinyLM(vocab_size, d, heads, layers, max_len), train_lm(model, seqs, epochs, lr, batch, seed), next_logprobs(model, prefixes); BM25(docs, k1, b), dense_scores, expand_query, hybrid; ToolRegistry(as_of) with register(name, fn, dated), call(name, **kw) and log.
Rules. No network calls and no downloaded weights in tests; scores reported by period relative to the training cut-off; every tool call logged; dated tools refuse anything after the as-of date.
Acceptance tests. code/firm/llmeval/tests/: the language model memorises a small corpus and is causal (a token’s prediction does not change with later tokens); training is deterministic; BM25 matches its formula on a hand example; query expansion adds only close neighbours; the registry refuses a future document, logs every call and rejects unknown tools.
Stretch. A local open-weights model behind the same interface, run offline; an entity-swap test; a prompt and version registry in the experiment tracker (chapter 25).
Sources and further reading
- A. Vaswani and co-authors, “Attention is all you need”, arXiv:1706.03762, 2017.
- P. Lewis and co-authors, “Retrieval-augmented generation for knowledge-intensive NLP tasks”, arXiv:2005.11401, 2020.
- S. Robertson and H. Zaragoza, “The probabilistic relevance framework: BM25 and beyond”, Foundations and Trends in Information Retrieval 3(4), 2009.
- N. Carlini and co-authors, “Extracting training data from large language models”, arXiv:2012.07805, 2021.
- A. Lopez-Lira and Y. Tang, “Can ChatGPT forecast stock price movements? Return predictability and large language models”, arXiv:2304.07619, 2023.
- P. Glasserman and C. Lin, “Assessing look-ahead bias in stock return predictions generated by GPT sentiment analysis”, arXiv:2309.17322, 2023.
- S. Wu and co-authors, “BloombergGPT: a large language model for finance”, arXiv:2303.17564, 2023.
- S. Yao and co-authors, “ReAct: synergizing reasoning and acting in language models”, arXiv:2210.03629, 2023.
- A. Kim, M. Muhn and V. Nikolaev, “Financial statement analysis with large language models”, arXiv:2407.17866, 2024.
14.8 Exercises
Exercise 14.1 ★
A model’s documentation gives a training cut-off of 30 June. A backtest runs from January of the previous year to December. Which part of it can measure the model’s judgement, and what does the rest measure?
Solution
Solution of Exercise 14.1.
Only July to December of the second year, after the cut-off, can measure judgement, and only if the provider’s cut-off is right and the model was not updated since. The eighteen months before it measure a mixture of judgement and memory of what followed each headline.
Exercise 14.2 ★
Compute the BM25 contribution of a term with document frequency 45 in 540 passages, appearing once in a passage of average length, with and .
Solution
Solution of Exercise 14.2.
. At average length the normaliser is , so the term factor is and the contribution is 2.48.
Exercise 14.3 ★
From Figure 14.2, compute each model’s memorisation premium and anonymisation loss.
Solution
Solution of Exercise 14.3.
Premium (between minus after): contaminated points, clean . Anonymisation loss between the cut-offs: contaminated , clean .
Exercise 14.4 ★★
Why does dense retrieval alone find almost none of the right passages, and why does adding its score to BM25’s help so little?
Solution
Solution of Exercise 14.4.
Its score depends only on the words it has embeddings for, the metric synonyms and context words; company names and years are not in its vocabulary, so the 135 passages about the right metric tie and the top one is about the right company by chance (1.1%). Added to BM25, it lifts every passage about the right metric equally, including the other company’s passage that shares the question’s synonym, which BM25 already ranks first: the two scores disagree only where dense retrieval is blind.
Exercise 14.5 ★★
Why is the contaminated model better than the clean one after both cut-offs (59.2% against 56.3%)? Is that evidence of judgement?
Solution
Solution of Exercise 14.5.
It trained on a third more headlines (7 988 to day 1 210 instead of 5 972 to day 907), so it estimates the events’ effects better, and more recently. After both cut-offs it has seen none of the answers, so the gain is generalisation, judgement in the chapter’s sense, bought with data. The point of the protocol is that this comparison is fair and the one between the cut-offs is not.
Exercise 14.6 ★★
Find the flaw. “We asked the model for the direction of 10 000 headlines from 2015 to 2023 and it was right 70% of the time, far above the 56% of our tf-idf model; we removed the company names as a check and it was still 65%.”
Solution
Solution of Exercise 14.6.
Every headline is before the model’s cut-off, so the 70% mixes memory and judgement; the comparison with a tf-idf model trained and tested out of sample is unfair. Removing the company names is not a full anonymisation (dates, products, people and unique phrasing identify the event), and 65% after it only shows that some memory survived or that the text is informative, not which. Score on headlines after the cut-off, anonymise names and dates, and compare with baselines on the same headlines.
Exercise 14.7 ★★★
Coding. Verify that TinyLM is causal: change the last token of a prompt and check that the model’s predictions at all earlier positions are unchanged. What in the code guarantees it?
Solution
Solution of Exercise 14.7.
Changing the last token leaves the logits at every earlier position unchanged (the acceptance test checks it to ). The attention mask is upper-triangular with above the diagonal, so after the softmax position puts zero weight on positions after ; the feed-forward layers act position by position, so nothing else mixes positions.
Exercise 14.8 ★★★
Design an agent task (research question, tools, as-of date, checks) for “which of our signals decayed most last quarter?”, and list what the log must contain for a reviewer to reproduce the answer.
Solution
Solution of Exercise 14.8.
Question: the signals whose information coefficient fell most between the previous and the last quarter, with the decay measure of Book 7 (chapter 13). Tools: the signal registry, the point-in-time data store with the as-of date set to the quarter end, the IC calculator, and a plotting tool; no tool may write. Checks: the as-of guard, the minimum number of observations per signal, a multiple-testing adjustment for ranking many signals. The log: task, model version, prompt, every tool call with arguments and results’ hashes, refusals, the final answer, and the code that recomputes it without the agent.
14.9 Problem: The Model That Knew
Problem 14.1
Weekend problem — memory or judgement
The chapter’s two language models, its filings collection and its agent.
Part I — The models.
- What does each model see in training, and how is it asked about a headline?
- How large is the model and its vocabulary, and why does the protocol matter more than the size?
- What is the best accuracy available after both cut-offs, and why is it not higher?
- Why is one training headline in ten anonymised?
Part II — Contamination.
- What are the two models’ accuracies before both cut-offs, and what does the anonymised accuracy there show?
- What are they between the cut-offs, and what explains the difference?
- What does anonymisation do to each model between the cut-offs?
- What does the literature say about the same two effects in real models?
Part III — Retrieval and agents.
- How do the four searches do with the question in the passage’s words and in other words?
- What does the grounding check cost and what does it buy?
- What does the agent do without its guard?
- What should the log of an agent’s run contain?
Part IV — The verdict.
- State the named result: the memorisation premium and the anonymisation loss for the contaminated and the clean model.
- How would you evaluate a vendor’s news-sentiment model built on an LLM?
- Which tasks would you give a language model first on a research desk?
- What must never go into a prompt sent outside the firm?
- Why can an agent’s backtest be worse than a person’s?
- How do you make an LLM-based research result reproducible?
- When would you fine-tune a model rather than prompt it?
- In one sentence: what does a language model’s accuracy on history measure?
Solution
Solution of Problem 14.1.
Part I.
- Sentences of a week token, the company, the headline and the outcome after an arrow, for every headline before its cut-off; asked by prompting up to the arrow and comparing the probabilities of “up” and “down”.
- 106 681 parameters and 505 tokens. Whether a model memorises depends on how often it saw a document and on what keys identify it, not only on size; the protocol (period split, anonymisation) measures memory in any model.
- 61.6%: the sign of the planted effect is right 66.0% of the time on the 1 462 headlines with an effect, and the other 550 are coin flips. The reaction’s noise (3%) is large against the effects.
- So that the placeholders are words the model has seen; otherwise the anonymisation test would measure the model’s reaction to unknown tokens, not the loss of its memory.
Part II.
- Clean 69.7%, contaminated 65.4%; anonymised 57.7% and 57.2%. Both memorised the outcomes of their training headlines, keyed by week and company; without the keys they fall to their judgement.
- Contaminated 64.8%, clean 56.2%: the contaminated model saw these headlines and their outcomes.
- It takes the contaminated model to 57.0% (a loss of 7.8 points) and the clean one to 56.0% (0.2).
- Lopez-Lira and Tang evaluated on headlines after GPT-4’s cut-off; Glasserman and Lin found that anonymised headlines did better in-sample, because general knowledge of the company distracted the model more than look-ahead helped it.
Part III.
- Same word: BM25 100%, dense 0%, hybrid 100%, expanded BM25 86.3%. Other word: 5.0%, 1.1%, 18.2%, 89.1%.
- It turns away 11.9% of the answerable questions, exactly those whose top passage named another company or year and would have been answered wrongly; it abstains on all 300 unanswerable questions, which would otherwise all get a number.
- It answers 80.8% of the as-of questions from a filing published later, and is right 19.2% of the time; with the guard 100%.
- Each tool call with its arguments, the as-of date, the result or refusal, and the model version and prompt.
Part IV.
- The model that knew. Memorisation premium 5.6 points for the contaminated model and for the clean one; anonymisation loss 7.8 points and 0.2.
- Ask for its cut-off and version history, results by period relative to the cut-off, an anonymised run, and a live or post-cut-off trial against a dictionary baseline.
- Extraction, tagging and summarisation with the source quoted, and code drafting with tests: outputs a person or a test can check.
- Material non-public information, client data and anything the firm would not publish.
- It runs more experiments, faster, without memory of what it tried, and uses whatever data the tools return; without a guard and a log it overfits and looks ahead at scale.
- Pin the model version, prompt, decoding settings and retrieval index; log every call; keep a code path that recomputes the result without the model.
- When the task is stable, labelled data exist, and prompting leaves accuracy or cost on the table; and when the model can be hosted where the data may go.
- Before its cut-off, memory and judgement together; only after it, judgement.
14.10 Interview questions
Interview question 14.1 ★ researcher, mle
What is a training cut-off, and why does it matter for a backtest that uses an LLM?
Solution
Solution of Interview question 14.1.
The date of the last text in the training data. Before it the model may have read the event and its aftermath, so a backtest over that period has look-ahead bias that no data hygiene on the firm’s side can remove.
What the interviewer is looking for: look-ahead through the model’s training data, and the remedy (post-cut-off evaluation).
Interview question 14.2 ★★ mle
Explain retrieval-augmented generation. How would you evaluate the retrieval step separately from the generation?
Solution
Solution of Interview question 14.2.
Retrieve passages, put them in the prompt, generate an answer grounded in them. Evaluate retrieval with labelled question–passage pairs (recall at ), and generation given the right passage (faithfulness, exact match), so that failures are attributed to the right stage.
What the interviewer is looking for: the two stages and separate metrics for each.
Interview question 14.3 ★★ researcher
How would you test whether an LLM’s apparent forecasting skill is memorisation?
Solution
Solution of Interview question 14.3.
Compare accuracy before and after the stated cut-off; anonymise names and dates and measure the loss; test on events the model cannot know (post-cut-off, or synthetic ones); check whether it can complete the headline’s text verbatim.
What the interviewer is looking for: the period split and anonymisation, ideally a completion test.
Interview question 14.4 ★★ mle, developer
BM25 or embeddings for searching filings? Defend a design.
Solution
Solution of Interview question 14.4.
BM25 for names, tickers, numbers and exact terms; embeddings for paraphrase; in practice both, with the embedding used to expand the query or rerank BM25’s candidates, measured on labelled questions from the desk’s own documents.
What the interviewer is looking for: what each misses, and a measured hybrid.
Interview question 14.5 ★★ mle
How do you stop a research agent from using data after the as-of date, and how would you prove it did not?
Solution
Solution of Interview question 14.5.
Put the as-of date in the harness, not the prompt: every data tool filters by publication time and refuses later results, and every call is logged. Prove it by the log and by a test that plants future documents and checks they are refused.
What the interviewer is looking for: enforcement in the tools, a log, and a planted-future test.
Interview question 14.6 ★★★ researcher, trader
A vendor sells an LLM sentiment score with a backtested Sharpe ratio of 3 over ten years. What do you ask for before a trial?
Solution
Solution of Interview question 14.6.
The model and version behind it and their cut-offs; results split before and after the cut-off; an anonymised backtest; the live history since launch; the timestamps of scores against news; costs and capacity; and a free trial on data after the cut-off.
What the interviewer is looking for: the cut-off split, anonymisation, live history and timestamps.