Tokenisation splits a text into units (words, word pieces, punctuation), usually after normalising case. A bag of words represents a document by the counts of its tokens, forgetting their order. An n-gram is a sequence of consecutive tokens; bags of n-grams keep some order. Term frequency–inverse document frequency (tf-idf) weights the count of term in document by , where of the documents contain , so that words found everywhere count for little; the rows are then scaled to unit length.
Quantitative Finance · Glossary
What is Tokenisation, bag of words, n-gram, tf-idf?
Also known as: tokenisation · bag of words · n-gram · term frequency--inverse document frequency