Traffic violations continue to be a major challenge for road safety and urban mobility. This has created growing interest in predictive analytics and intelligent transportation systems. As cities collect more traffic data, machine learning has become a useful tool for predicting violations and supporting proactive enforcement. This thesis presents a data-driven framework for predicting traffic violations using a large dataset that covers one full year of recorded incidents. The dataset includes important temporal, spatial, and vehicle-related features that provide a good foundation for building classification models.
A systematic data preparation process was developed. This included data cleaning, numerical encoding, min-max normalization, and refinement of violation categories to focus on the most common patterns. To deal with the serious class imbalance among violation types, several resampling techniques were tested. Random oversampling gave the best improvement in model performance. These steps produced a balanced and consistent dataset suitable for machine learning.
One of the main contributions of this thesis is an improved Random Forest model. Two important changes were made to enhance its performance. First, the standard Gini impurity measure was replaced with Information Gain to create more meaningful feature splits. Second, Out-of-Bag (OOB) error estimation was integrated to provide unbiased performance evaluation during training. Moreover, this kind of research is considered as first of a kind in Qatar and in the region which support Qatar National Vision 2030 under the Safety aspects and road safety strategy. The proposed model was compared with several established algorithms: Naïve Bayes, Logistic Regression, Support Vector Machine, XGBoost, Neural Networks, and the standard Random Forest. The results show that the improved Random Forest achieved 99.5% accuracy and performed very well across precision, recall, and F1-score.
The thesis also proposes a new deep learning classification framework, TrafficViolationNet, belonging to the class of residual-attention neural architectures, which combines an enhanced residual architecture with a lightweight attention mechanism to improve feature learning from structured traffic data.
Feature importance analysis revealed that street number is the most influential predictor. This highlights the strong role of spatial patterns in driver behaviour. In contrast, license plate type had very little influence. The final model can generate scenario-based predictions and spatial visualizations. These outputs help identify high-risk zones and streets and provide practical insights for traffic authorities.
Overall, this thesis presents an effective, scalable, and interpretable machine learning framework for traffic violation prediction. By combining careful data preprocessing with an enhanced Random Forest model, the study demonstrates the value of predictive analytics in improving road safety and supporting intelligent traffic management systems.
| Date of Award | 2026 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - HBKU College of Science and Engineering
|
|---|
- Data Analysis
- Machine Learning
- Prediction
- Random Forest
- Safety
- Traffic Violation
ARTIFICIAL INTELLIGENCE AND ITS IMPLEMENTATION IN IDENTIFYING AND PREDICTING TRAFFIC VIOLATIONS
Alshriem, M. (Author). 2026
Student thesis: Doctoral Dissertation