A topic model describes each document as a mixture of a few topics, each topic a distribution over words, all estimated from word counts without labels. Latent Dirichlet allocation (LDA) is the Bayesian topic model in which each document’s topic shares and each topic’s word probabilities have Dirichlet priors, and each word is drawn by first drawing its topic from the document’s shares (Blei, Ng and Jordan, 2002).
| topic | six most probable words |
|---|---|
| 1 | analysts, fell, short, expectations, forecasts, exceeded |
| 2 | outlook, guidance, but, expectations, forecasts, estimates |
| 3 | annual, meeting, record, attendance, reports, holds |
| 4 | full, forecast, year, reports, line, costs |
| 5 | wins, authorises, program, repurchase, industry, award |
| 6 | announces, gets, rating, earnings, schedules, call |
ml_text.topics.