In the previous post, we learned Word Embedding. In this post, we learn Sentence Embedding and Document Embedding.
In the previous post, we learned word embeddings i.e. for the smallest units in language word. However, in most business applications dealing with NLP, understanding longer units of meaning such as sentences and documents is crucial. In this post, we learn sentence and document embeddings.
Introduction to Sentence and Document Embedding
Sentence and Word embeddings often are derived from word embeddings.
There are various techniques to generate sentence and document embeddings.
Technique 1 (using BERT):
Use BERT for sentence representations
Feed the entire document into BERT.
Technique 2 (Average of each word embeddings across the documents):
In a sentence or a document, firstly create each word embeddings a floating-point number, secondly add up each word embedding, and finally take the average of this to generate the document word embeddings.
Technique 3 (Max Pooling i.e. maximum of the numbers instead of the Average):
In a sentence or a document, first, create each word embeddings a floating-point number, second add up each word embedding, and finally take the maximum of the numbers instead of the average. This is called the max-pooling technique.
Technique 4: (TF-IDF identifies the most important words for the topic)
Through TF-IDF you can determine the most important words for a topic. Through this technique, you can also determine which words are more or less likely to appear related to a topic.
Which technique to use for Sentence and Document Embedding
For Documents, use the max pool technique for document classification.
For Sentences, use LSTM kind of sequence representation methods. Or, use BERT sequence generation.
Related Topics:
- Natural Language Processing Primer
- NLP Learnings 02 – After Primer – NLP Models & Evaluating NLP Models
- NLP Learnings 03 – What is Linguistic in NLP and Linguistic Categories or Linguistic Levels
- NLP Learnings 04 – Sentence Linguistic Analysis – Words As First Step
- NLP Learnings 05 – Sentence Linguistic Analysis – Pragmatics Analysis
- NLP Learnings 06 – What can we do with A Sentence in NLP Tasks
- NLP Learnings 07 – Introducing NLP Language Models
- NLP Learnings 08 – Language Models – Probability Types
- NLP Learnings 09 – Language Models – Measuring Text Based On Probability
- NLP Learnings 10 – Language Models – Defining Word Boundaries
- NLP Learnings 11 – Language Models – Your Business Text Data and Their Words Representation
- NLP Learnings 12 – Language Models – Regularization Techniques And Your Business Text Data
- NLP Learnings 13 – Language Models – Smoothing Regularization Techniques And Your Business Text Data
- NLP Learnings 14 – Language Models – NLP Key Terms and Concepts
- NLP Learnings 15 – Language Models – Classifiers Introduction
- NLP Learnings 16 – Language Models – Classifiers – What are Probabilistic Classifiers
- NLP Learnings 17 – Language Models – Classifiers – Simple Classifiers Introduction
- NLP Learnings 18 – Language Models – Classifiers – Simple Classifiers – Linear Probabilistic Classifier Introduction
- NLP Learnings 19 – Language Models – Classifiers – Evaluating Classifiers – Precision Recall F-Score Confusion Matrix
- NLP Learnings 20 – Language Models – Classifiers – Represent Words In A Document – Choosing and Representing Features In The Right Way
- NLP Learnings 21 – Language Models – Word Embeddings
- NLP Learnings 22 – Language Models – Sentence and Document Embeddings
- NLP Learnings 23 – Language Models – Sentence and Document Embeddings – Which one to consider