Skip to main navigation Skip to search Skip to main content

A VIDEO ANALYTICS FRAMEWORK OF SOCCER GAMES USING AI

  • Fahad Majeed

Student thesis: Doctoral Dissertation

Abstract

This thesis presents a unified, real-time framework for advanced soccer video analytics, integrating object detection, instance segmentation, multi-object tracking, ball-player interaction modelling, and multimodal performance evaluation. To enhance conventional RGB-based analysis, motion vectors and frame differencing are incorporated, providing spatiotemporal cues for improved object discrimination under occlusion and high-speed dynamics. We first introduce MV-Soccer, based on CSPDarknet53 and GELAN backbones for joint segmentation and tracking. Trained on DFL-Bundesliga, SoccerNet-Tracking, and SoccerPro, it achieves up to 99% accuracy on SoccerPro, 98% on SoccerNet-Tracking, and 97% on DFL-Bundesliga. Building on this, ReST enhances real-time segmentation and tracking by extracting Scharr-filter-based motion vectors for improved robustness in dense scenarios. For higher-level tactical analysis, we propose GameFlow, a sequential fusion pipeline using Graph Convolutional Networks (GCNs) to model ball-player interactions. By encoding inter-player relationships and motion features such as speed and distance, GameFlow enables fine-grained strategy analysis, achieving 91% detection, 90% tracking, action recognition and 92% speed estimation performance. To support comprehensive evaluation, we introduce 3MT (Multimodal Multitask Learning), which jointly processes visual, textual, audio, and video modalities for detection, tracking, and captioning. The framework integrates SigLIP with state-space models in a Mamba-YOLO architecture, improving efficiency and cross-modal alignment. We also present 3MT++, a large-scale dataset with 19.9 M annotations across 9.4 M images. This approach yields a 5% improvement in player performance prediction and a 3% gain in decision-making accuracy. Finally, we present EchoNet, a multilingual audio analysis pipeline for soccer broadcasts, including audio extraction (FFmpeg), filtering (FFT with a Butterworth bandpass filter: 300–7 kHz), denoising (Demucs), segmentation (Silero VAD), and transcription (Whisper variants). Outputs are translated and structured into match-level JSON. Evaluation shows low word error rates and robust performance across diverse conditions. Overall, this thesis delivers a scalable framework combining visual analytics, graph-based modelling, multimodal learning, and audio understanding to provide comprehensive insights into player performance, decision-making, and team strategies.
Date of Award2026
Original languageAmerican English
Awarding Institution
  • HBKU College of Science and Engineering

Keywords

  • Computer Vision
  • Decision Making
  • Deep Learning
  • Multimodal Learning
  • Sports Video Analytics

Cite this

'