Tokenisation splits a text into units (words, word pieces, punctuation), usually after normalising case. A bag of words represents a document by the counts of its tokens, forgetting their order. An n-gram is a sequence of consecutive tokens; bags of n-grams keep some order. Term frequency–inverse document frequency (tf-idf) weights the count of term in document by , where of the documents contain , so that words found everywhere count for little; the rows are then scaled to unit length.
Quantitative Finance · Glossaire
Qu'est-ce que « Tokenisation, bag of words, n-gram, tf-idf » ?
Aussi appelé : tokenisation · bag of words · n-gram · term frequency--inverse document frequency