Sentiment Analysis
Sentiment Analysis is one of the most classic and widely applied tasks in the field of Natural Language Processing (NLP). It uses computational techniques to automatically identify, extract, and analyze subjective information in text, determining whether the author's attitude toward a specific topic, product, or service is positive, negative, or neutral.
Basic Types of Sentiment Analysis
Classification by Analysis Granularity
- Document-Level Sentiment Analysis: Determines the sentiment tendency of the entire document as a whole
- Sentence-Level Sentiment Analysis: Analyzes the sentiment polarity of a single sentence
- Aspect-Level Sentiment Analysis: Makes sentiment judgments on specific aspects mentioned in the text
Classification by Sentiment Dimension
- Binary Classification: Positive/Negative
- Three-Class Classification: Positive/Neutral/Negative
- Multi-Class Classification: More fine-grained sentiment classification (such as anger, happiness, sadness, etc.)
- Sentiment Intensity Analysis: Quantifies the intensity of sentiment
Dictionary-Based Sentiment Analysis Method
The dictionary-based method is the most traditional sentiment analysis technique, relying mainly on pre-built sentiment dictionaries.
Core Components
Sentiment Dictionary: A collection of words with sentiment polarity and intensity
- Common English dictionaries: SentiWordNet, AFINN, VADER
- Common Chinese dictionaries: HowNet Sentiment Dictionary, Dalian University of Technology Sentiment Vocabulary Ontology Database
Intensity Modifier: Handles the influence of degree adverbs and negation words
- Degree adverbs: very (1.5), quite (1.3), a bit (0.8), etc.
- Negation words: not, no, absolutely not, etc.
Basic Workflow
Example
def lexicon_based_sentiment(text):
sentiment_score = 0
words = tokenize(text) # Word segmentation
for word in words:
if word in positive_lexicon:
sentiment_score += positive_lexicon[word]
elif word in negative_lexicon:
sentiment_score -= negative_lexicon[word]
# Handle negation and degree modification
sentiment_score = apply_negation(words, sentiment_score)
sentiment_score = apply_intensifier(words, sentiment_score)
return normalize(sentiment_score)
Advantages and Disadvantages Analysis
Advantages:
- No training data required
- High computational efficiency
- Strong interpretability
Disadvantages:
- Difficult to handle complex linguistic phenomena (such as sarcasm, irony)
- Relies on the coverage and quality of the dictionary
- Cannot capture contextual semantics
Machine Learning-Based Sentiment Analysis Method
Machine learning methods perform sentiment analysis by learning patterns from annotated data.
Typical Feature Engineering
- Bag of Words Model (BOW): Represents text as a vector of word occurrence frequencies
- TF-IDF: Considers the importance of words in the document
- N-gram Features: Captures local word sequence patterns
- Sentiment Dictionary Features: Combines the advantages of dictionary methods
Common Algorithms

Code Example: Sentiment Classification with Scikit-learn
Example
from sklearn.svm import LinearSVC
from sklearn.pipeline import Pipeline
# Build classification pipeline
sentiment_clf = Pipeline([
('tfidf', TfidfVectorizer(ngram_range=(1, 2))),
('clf', LinearSVC())
])
# Train the model
sentiment_clf.fit(train_texts, train_labels)
# Predict new text
prediction = sentiment_clf.predict(["This product is extremely easy to use, highly recommended!"])
print(prediction) # Output: 'positive'
Fine-Grained Sentiment Analysis
Aspect-Based Sentiment Analysis (ABSA) is a more advanced sentiment analysis task that aims to identify specific aspects mentioned in text and their corresponding sentiments.
Core Subtasks of ABSA
Aspect Extraction: Identifies entities or attributes discussed in the text
- Explicit aspect: "The phone's battery life is very good" → "battery"
- Implicit aspect: "The photos taken are very clear" → "camera"
Sentiment Classification: Makes sentiment judgments for each identified aspect
Comparison of Implementation Methods
| Method Type | Representative Model | Applicable Scenario | Advantages | Disadvantages |
|---|---|---|---|---|
| Pipeline Method | First use CRF to extract aspects, then use a classifier to determine sentiment | Scenarios with limited resources | Clear modules, easy to debug | Error propagation |
| End-to-End Method | BERT-ABSA、AOA-LSTM | High precision requirements | Joint optimization, better performance | Requires more data |
| Multi-Task Learning | MT-DNN、Multi-Task BERT | Related task assistance | Knowledge sharing | Difficulty in task balancing |
Code Example: BERT-Based Aspect-Level Sentiment Analysis
Example
import torch
# Load pre-trained model
model = BertForSequenceClassification.from_pretrained('bert-base-uncased', num_labels=3)
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
# Prepare input
text = "The restaurant's environment is great, but the service is too slow."
aspect = "service"
inputs = tokenizer(f"[CLS] {aspect} [SEP] {text} [SEP]", return_tensors="pt")
# Predict sentiment
outputs = model(**inputs)
predictions = torch.argmax(outputs.logits, dim=1)
print(predictions) # Possible output: 1 (negative)
Challenges and Development Directions of Sentiment Analysis
Current Main Challenges
- Context Dependence: The same word may have different sentiments in different contexts
- Domain Adaptability: Models trained in one domain perform worse in other domains
- Multilingual Processing: Sentiment expression varies greatly across different languages
- Sarcasm and Irony Detection: Cases where the literal text is opposite to the actual sentiment
Frontier Development Directions
- Multimodal Sentiment Analysis: Combines multiple types of information such as text, images, and speech
- Cross-Lingual Sentiment Analysis: Leverages commonalities between languages to improve performance on low-resource languages
- Sentiment Cause Extraction: Not only determines sentiment, but also analyzes the causes
- Personalized Sentiment Analysis: Considers the user's personal characteristics and historical behavior
Practical Exercises
Exercise 1: Building a Basic Sentiment Analyzer
- Implement a simple sentiment analyzer using NLTK's VADER dictionary
- Test its accuracy on a movie review dataset
Exercise 2: Comparing Different Machine Learning Methods
- Train sentiment classifiers using Naive Bayes, SVM, and Logistic Regression respectively
- Use cross-validation to compare their performance differences
Exercise 3: Aspect-Level Sentiment Analysis Practice
- Fine-tune a pre-trained BERT model on the SemEval 2014 restaurant review dataset
- Implement an end-to-end system that can simultaneously extract aspects and determine sentiment
Through this article, you should have mastered the basic concepts, main methods, and implementation techniques of sentiment analysis. As a fundamental NLP task, sentiment analysis technology continues to evolve and has broad value in practical applications, playing an important role from product review analysis to social media monitoring.
Other Extensions