This project presents a comprehensive comparative study of deep learning and transformer-based architectures for Twitter sentiment classification.
Four different models were implemented and evaluated on a large-scale Twitter sentiment dataset containing over 1 million tweets:
- πΉ Simple Recurrent Neural Network (RNN)
- πΉ Long Short-Term Memory Network (LSTM)
- πΉ Gated Recurrent Unit (GRU)
- πΉ Bidirectional Encoder Representations from Transformers (BERT)
The objective is to analyze how modern neural architectures improve sentiment classification performance and identify the most suitable model for large-scale sentiment analysis.
| Attribute | Value |
|---|---|
| Total Tweets | 1,048,572 |
| Features | 6 |
| Negative Tweets | 799,996 |
| Positive Tweets | 248,576 |
| Sentiment Classes | Binary Classification |
Distribution of positive and negative sentiment classes.
βββββββββββββββββββββββββββββββββββββββββββββββ
β Twitter Sentiment Dataset β
β 1,048,572 Tweets β
βββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββ
β Exploratory Data Analysis β
β β’ Class Distribution β
β β’ Tweet Length Analysis β
β β’ Word Frequency Analysis β
β β’ Positive/Negative Word Clouds β
βββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββ
β Text Preprocessing β
β β’ Lowercasing β
β β’ URL & Username Removal β
β β’ Punctuation Removal β
β β’ Whitespace Normalization β
βββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββ
β Train / Validation / Test β
β 175k / 37.5k / 37.5k Samples β
βββββββββββββββββββββββββββββββββββββββββββββββ
β
βββββββββββββΌββββββββββββ
β β β
βΌ βΌ βΌ
βββββββββββββββ βββββββββββββββ βββββββββββββββ
β Vocabulary β β Vocabulary β β Vocabulary β
β 20,000 β β 20,000 β β 20,000 β
β Max Len=32 β β Max Len=32 β β Max Len=32 β
ββββββββ¬βββββββ ββββββββ¬βββββββ ββββββββ¬βββββββ
β β β
βΌ βΌ βΌ
βββββββββββββββ βββββββββββββββ βββββββββββββββ
β RNN β β LSTM β β GRU β
β Train from β β Train from β β Train from β
β Scratch β β Scratch β β Scratch β
ββββββββ¬βββββββ ββββββββ¬βββββββ ββββββββ¬βββββββ
β β β
βββββββββ¬ββββββββ΄ββββββββ¬ββββββββ
β β
βΌ β
βββββββββββββββββββ β
β Performance β β
β Evaluation β β
ββββββββββ¬βββββββββ β
β β
βΌ β
βββββββββββββββ β
β Accuracy β β
β Precision β β
β Recall β β
β F1 Score β β
β ROC-AUC β β
βββββββββββββββ β
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββ
β BERT β
β bert-base-uncased β
β Pretrained Transformer β
β Full Fine-Tuning β
β Max Length = 48 β
βββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββ
β Comparative Analysis β
β RNN vs LSTM vs GRU vs BERT β
β Accuracy β’ Precision β’ Recall β
β F1 Score β’ ROC-AUC β
βββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββ
β Best Model: BERT β
β Accuracy : 87.72% β
β F1 Score : 72.48% β
β ROC-AUC : 92.21% β
βββββββββββββββββββββββββββββββββββββββββββββββ
The project follows a complete NLP pipeline:
- Exploratory Data Analysis
- Text Preprocessing
- Dataset Splitting
- Sequence Preparation
- Model Development
- Performance Evaluation
- Comparative Analysis
- Best Model Selection
Key findings from the dataset:
- No missing values detected
- Strong class imbalance (76% Negative, 24% Positive)
- Average tweet length β 74 characters
- Median tweet length β 70 characters
- Significant vocabulary overlap between sentiment classes
Most frequently occurring words in positive and negative tweets.
The following preprocessing pipeline was applied:
- Lowercasing
- URL Removal
- Username Removal
- Punctuation Removal
- Numeric Filtering
- Whitespace Normalization
- Vocabulary Size: 20,000
- Maximum Sequence Length: 32
- Pretrained BERT Tokenizer
- Maximum Sequence Length: 48
- Attention Masks
- Special Tokens
Baseline recurrent neural network trained from scratch.
- Embedding Layer (128)
- Simple RNN Layer (128 Hidden Units)
- Dropout (0.3)
- Dense Output Layer
Training and validation loss during RNN training.
Long Short-Term Memory architecture designed to capture long-range contextual dependencies.
- Embedding Layer (128)
- LSTM Layer (128 Hidden Units)
- Dropout (0.3)
- Dense Output Layer
Training and validation loss during LSTM training.
Gated Recurrent Unit architecture providing efficient sequence modeling with fewer parameters.
- Embedding Layer (128)
- GRU Layer (128 Hidden Units)
- Dropout (0.3)
- Dense Output Layer
Training and validation loss during GRU training.
Pretrained transformer model fine-tuned end-to-end for sentiment classification.
- BERT Base (Uncased)
- Transformer Encoder
- Classification Head
- Full Fine-Tuning
- AdamW Optimizer
- Learning Rate = 2e-5
- Best Model Checkpointing
Training and validation loss during BERT fine-tuning.
| Model | Accuracy | Precision | Recall | F1 Score | ROC-AUC |
|---|---|---|---|---|---|
| RNN | 0.7530 | 0.4820 | 0.5637 | 0.5197 | 0.7084 |
| LSTM | 0.8117 | 0.5789 | 0.7541 | 0.6550 | 0.8723 |
| GRU | 0.8030 | 0.5613 | 0.7730 | 0.6504 | 0.8697 |
| BERT | 0.8772 | 0.7730 | 0.6822 | 0.7248 | 0.9221 |
Comparison of Accuracy, Precision, Recall, F1 Score, and ROC-AUC across all architectures.
- Established baseline performance
- Limited contextual understanding
- Lowest overall performance
- Significant improvement over vanilla RNN
- Strong sequence modeling capability
- Approximately 26% improvement in F1 Score
- Comparable performance to LSTM
- Faster and computationally efficient
- Highest Recall score
- Best overall performance
- Highest Accuracy, Precision, F1 Score, and ROC-AUC
- Benefited from pretrained contextual language representations
| Metric | Score |
|---|---|
| Accuracy | 87.72% |
| Precision | 77.30% |
| Recall | 68.22% |
| F1 Score | 72.48% |
| ROC-AUC | 92.21% |
sentiment-analysis-comparison/
βββ data/
βββ notebooks/
β βββ 01_EDA.ipynb
β βββ 02_Preprocessing.ipynb
β βββ 03_RNN.ipynb
β βββ 04_LSTM.ipynb
β βββ 05_GRU.ipynb
β βββ 06_BERT.ipynb
β βββ 07_Comparative_Analysis.ipynb
β
βββ models/
βββ images/
βββ reports/
βββ results/
β
βββ README.md
βββ requirements.txt
- Python
- Pandas
- NumPy
- Matplotlib
- Seaborn
- WordCloud
- PyTorch
- Hugging Face Transformers
- Scikit-Learn
- Hyperparameter Optimization
- Threshold Tuning
- RoBERTa Comparison
- DeBERTa Comparison
- Multi-Class Sentiment Classification
- Real-Time Sentiment Monitoring Dashboard
- FastAPI Deployment
- Streamlit Application
This project demonstrates the progression of sentiment classification performance from traditional recurrent neural networks to transformer-based language models.
While LSTM and GRU significantly outperformed the vanilla RNN baseline, BERT achieved the strongest overall performance with an Accuracy of 87.72%, an F1 Score of 72.48%, and a ROC-AUC of 92.21%.
The results highlight the effectiveness of transfer learning and contextual language understanding in modern Natural Language Processing applications.










