TensorFlow Example - Text Classification Project
Text classification is a fundamental task in Natural Language Processing (NLP), which refers to automatically classifying text documents into one or more predefined categories. In practical applications, text classification is widely used for:
- Spam detection
- Sentiment analysis
- News classification
- Customer service conversation classification
- Product review classification
Implementing text classification with TensorFlow usually involves the following steps:
- Data preparation and preprocessing
- Text vectorization
- Model building
- Model training
- Model evaluation
- Model deployment
Environment Preparation
Before starting the project, make sure the following Python libraries are installed:
!pip install tensorflow !pip install numpy !pip install pandas !pip install matplotlib
Import the necessary libraries:
Example
from tensorflow.keras import layers
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
Check TensorFlow version:
Example
# Example output: 2.8.0
Dataset Preparation
We will use the IMDB movie review dataset, which is a classic binary classification dataset containing 50,000 movie reviews labeled as positive (1) or negative (0).
Load Dataset
Example
imdb = tf.keras.datasets.imdb
# Keep only the top 10,000 most frequently occurring words
(train_data, train_labels), (test_data, test_labels) = imdb.load_data(num_words=10000)
Data Exploration
View data format:
Example
# Output: Training samples: 25000, Test samples: 25000
# View the first review
print(train_data[0])
# Output: [1, 14, 22, 16, 43, 530, 973, 1622, 1385, 65, ...]
Data Preprocessing
Convert integer sequences to multi-hot encoding:
Example
results = np.zeros((len(sequences), dimension))
for i, sequence in enumerate(sequences):
results[i, sequence] = 1.
return results
x_train = vectorize_sequences(train_data)
x_test = vectorize_sequences(test_data)
# Convert labels to floating point numbers
y_train = np.asarray(train_labels).astype('float32')
y_test = np.asarray(test_labels).astype('float32')
Build Model
Model Architecture
We will build a simple fully connected neural network:
Example
layers.Dense(16, activation='relu', input_shape=(10000,)),
layers.Dense(16, activation='relu'),
layers.Dense(1, activation='sigmoid')
])
Model Compilation
Example
loss='binary_crossentropy',
metrics=['accuracy'])
Parameter description:
optimizer: Optimizer, controls the learning processloss: Loss function, measures the difference between model predictions and true labelsmetrics: Evaluation metric, monitors training and testing steps
Train Model
Create Validation Set
Example
partial_x_train = x_train[10000:]
y_val = y_train[:10000]
partial_y_train = y_train[10000:]
Training Process
Example
partial_y_train,
epochs=20,
batch_size=512,
validation_data=(x_val, y_val))
Visualize Training Results
Example
# Plot training loss and validation loss
plt.plot(history_dict['loss'], 'bo', label='Training loss')
plt.plot(history_dict['val_loss'], 'b', label='Validation loss')
plt.title('Training and validation loss')
plt.xlabel('Epochs')
plt.ylabel('Loss')
plt.legend()
plt.show()
# Plot training accuracy and validation accuracy
plt.plot(history_dict['accuracy'], 'bo', label='Training acc')
plt.plot(history_dict['val_accuracy'], 'b', label='Validation acc')
plt.title('Training and validation accuracy')
plt.xlabel('Epochs')
plt.ylabel('Accuracy')
plt.legend()
plt.show()
Model Evaluation and Prediction
Evaluate Test Set Performance
Example
print(results)
# Example output: [0.3245, 0.8732] represents loss and accuracy
Make Predictions
Example
print(predictions[0]) # Prediction probability of the first test sample
Model Optimization Suggestions
Adjust network architecture:
- Increase or decrease the number of hidden layers
- Try different numbers of neurons
- Use different activation functions
Regularization techniques:
- Add Dropout layers to prevent overfitting
- Use L1/L2 regularization
Optimizer selection:
- Try other optimizers such as Adam, SGD, etc.
- Adjust the learning rate
Text preprocessing improvements:
- Use word embeddings instead of multi-hot encoding
- Try pretrained word vectors (such as Word2Vec, GloVe)
Complete Code Example
Example
from tensorflow.keras import layers
import numpy as np
import matplotlib.pyplot as plt
# Load data
imdb = tf.keras.datasets.imdb
(train_data, train_labels), (test_data, test_labels) = imdb.load_data(num_words=10000)
# Data preprocessing
def vectorize_sequences(sequences, dimension=10000):
results = np.zeros((len(sequences), dimension))
for i, sequence in enumerate(sequences):
results[i, sequence] = 1.
return results
x_train = vectorize_sequences(train_data)
x_test = vectorize_sequences(test_data)
y_train = np.asarray(train_labels).astype('float32')
y_test = np.asarray(test_labels).astype('float32')
# Build model
model = tf.keras.Sequential([
layers.Dense(16, activation='relu', input_shape=(10000,)),
layers.Dense(16, activation='relu'),
layers.Dense(1, activation='sigmoid')
])
# Compile model
model.compile(optimizer='rmsprop',
loss='binary_crossentropy',
metrics=['accuracy'])
# Train model
history = model.fit(x_train, y_train,
epochs=4,
batch_size=512,
validation_data=(x_test, y_test))
# Evaluate model
results = model.evaluate(x_test, y_test)
print("Test loss and accuracy:", results)
# Make predictions
predictions = model.predict(x_test)
print("Prediction probability of the first review:", predictions[0])