- Project number
- Project / 003
- Year
- 2025
- Category
- Research
- Status
- Completed
Duplicate Question Detection (NLP)
My Master's research: can duplicate questions be detected accurately without the heavy models? FastText embeddings plus PCA kept 83.68% accuracy while cutting training time by two thirds and memory by 83%.
- Context
- Master of Science in AI, UMPSA (Master Project II)
- Data
- Quora Question Pairs · 404,290 pairs
- Stack
- Python · FastText · PCA · scikit-learn · XGBoost
- Result
- 83.68% accuracy at 200 dimensions
01 /Overview
Accurate enough, and cheap enough to run.
Duplicate questions fragment answers and waste moderators' time on large Q&A platforms. The strongest published detectors use deep models like BERT and BiLSTMs, which are accurate but expensive to train and serve.
This research asked a narrower, practical question: how much accuracy do you actually lose if you compress the features aggressively, and what do you gain in speed and memory?
02 /The problem
High-dimensional vectors are slow at scale.
Manual detection does not scale
Semantically identical questions are phrased differently, so keyword matching misses them.
Heavy features, heavy costs
Pair features built from word embeddings reach 1,200 dimensions, which slows training and inference.
A gap in the literature
Most prior work reports accuracy and ignores computational efficiency entirely.
03 /The idea
Compress the meaning, keep the signal.
- Preprocess
- FastText embeddings
- Pair features
- PCA
- Classify
- Evaluate
Represent each question with FastText word vectors, combine the pair into features, then use Principal Component Analysis to shrink those features before classification, and measure both accuracy and cost.
04 /System architecture
From raw text to a duplicate score.
Step 1: SOURCE
Quora pairs
404,290 question pairs, 36.9% duplicates.
Step 2: PROCESS
Preprocess
Cleaning, tokenisation, lemmatisation; question words like how and why kept.
Step 3: MODEL
FastText
300-dimensional vectors, 94.3% vocabulary coverage.
Step 4: PROCESS
Pair features
Absolute differences and element-wise products: 1,200 features.
Step 5: PROCESS
PCA
Reduced to 50, 100, 150 and 200 dimensions.
Step 6: OUTPUT
Classifier
Logistic Regression, SVM, Random Forest, XGBoost.
05 /Data / intelligence
200 dimensions keep almost everything.
FastText separated the classes clearly: duplicate pairs averaged a similarity of 0.742 against 0.431 for non-duplicates. PCA at 200 dimensions kept 89.5% of the variance while shrinking the features by 83.3%.
| Model | Training 1200D | Training 200D | Saved | Inference 200D |
|---|---|---|---|---|
| Logistic Regression | 42.3s | 9.87s | 76.7% | 0.8ms |
| SVM | 156.7s | 58.34s | 62.8% | 1.9ms |
| Random Forest | 89.4s | 32.78s | 63.3% | 0.9ms |
| XGBoost | 67.8s | 24.56s | 63.8% | 1.1ms |
Memory for feature storage fell from 3.2 GB to 0.53 GB (83.4%) for every model.
06 /Build
Systematic, not lucky.
01Four dimension settings
50, 100, 150 and 200 dimensions, each tested against the full baseline.
02Four classifiers
Linear and ensemble models compared on the same features.
03Question-aware cleaning
Interrogative words were kept because they carry intent.
04Cost as a metric
Training time, memory and inference speed measured alongside accuracy.
07 /Outcome
0.33% less accurate, far cheaper.
- Accuracy (200D, XGBoost)
- 83.68%
- Full 1,200D: 84.01%
- F1-score (200D)
- 81.22%
- Training time saved
- 66.7%
- average across models
- Memory saved
- 83.4%
- 3.2 GB → 0.53 GB
Deep-learning approaches in the literature reach 91 to 96% accuracy, but at far higher computational cost. This framework trades a few points of accuracy for a system that is fast and light enough to run widely.
08 /What I learned
Measure the cost, not just the score.
01The trade-off curve is the result
Showing accuracy against dimensions said more than any single best number.
02Most meaning survives compression
Losing five sixths of the features cost a third of a percentage point.
03Efficiency belongs in the evaluation
A model that is too expensive to deploy has not really solved the problem.
09 /Next project
Sepang Vision LabRace-intelligence lab for Sepang International Circuit: circuit geometry, simulated telemetry, replay, weather scenarios and Monte Carlo strategy.
