In the previous post, we learned different types of probabilities in Language Models (i.e. Joint Probability and Conditional Probability). In this post, we learn using the probability technique how we measure a text.
Types of Probability-Based Language Models
Unigram, Bi-Gram, Tri-Gram, N-Gram, and Full History Models are some of the probability-based Language Models.
Using the above probability models, we can predict the occurrences of a word in a text or corpus.
Why not the Measure Probability of a “whole Sentence” and Why we Measure Probability of Words, Phrases within a Sentence?
Generally, we determine the probability of words in a sentence and/or phrases in a sentence. We generally do not take the whole sentence as “one unit” for probability determination because sentences “as a single unit” generally do not repeat in a text or corpus.
Unigram Language Models (context is not the focus here)
It takes one word at a time in a corpus and finds each word’s probability in the corpus.
Since these Unigram Language Models do not look at the previous word, they lack the “context” of the word within the sentence or corpus.
Full History Language Models
These models are designed to consider the whole sentence as “one unit” and determine the probability of the whole sentence.
As said above, since sentences as a unit do not repeat within a corpus, the challenge with the Full History Language Models is, we need a huge training dataset to train the Full History Language Models. Through this huge training dataset, there could be chances of repeating sentence(s) and having a better probability value for a sentence “as a whole”.
Related Topics:
- Natural Language Processing Primer
- NLP Learnings 02 – After Primer – NLP Models & Evaluating NLP Models
- NLP Learnings 03 – What is Linguistic in NLP and Linguistic Categories or Linguistic Levels
- NLP Learnings 04 – Sentence Linguistic Analysis – Words As First Step
- NLP Learnings 05 – Sentence Linguistic Analysis – Pragmatics Analysis
- NLP Learnings 06 – What can we do with A Sentence in NLP Tasks
- NLP Learnings 07 – Introducing NLP Language Models
- NLP Learnings 08 – Language Models – Probability Types
- NLP Learnings 09 – Language Models – Measuring Text Based On Probability
- NLP Learnings 10 – Language Models – Defining Word Boundaries
- NLP Learnings 11 – Language Models – Your Business Text Data and Their Words Representation
- NLP Learnings 12 – Language Models – Regularization Techniques And Your Business Text Data
- NLP Learnings 13 – Language Models – Smoothing Regularization Techniques And Your Business Text Data
- NLP Learnings 14 – Language Models – NLP Key Terms and Concepts
- NLP Learnings 15 – Language Models – Classifiers Introduction
- NLP Learnings 16 – Language Models – Classifiers – What are Probabilistic Classifiers
- NLP Learnings 17 – Language Models – Classifiers – Simple Classifiers Introduction
- NLP Learnings 18 – Language Models – Classifiers – Simple Classifiers – Linear Probabilistic Classifier Introduction
- NLP Learnings 19 – Language Models – Classifiers – Evaluating Classifiers – Precision Recall F-Score Confusion Matrix
- NLP Learnings 20 – Language Models – Classifiers – Represent Words In A Document – Choosing and Representing Features In The Right Way
- NLP Learnings 21 – Language Models – Word Embeddings
- NLP Learnings 22 – Language Models – Sentence and Document Embeddings
- NLP Learnings 23 – Language Models – Sentence and Document Embeddings – Which one to consider