Skip to main navigation Skip to search Skip to main content

A LOW-RANK VIDEO REPRESENTATION FRAMEWORK FOR HUMAN ACTION RECOGNITION

  • Lujayn Al-Amir

Student thesis: Master's Dissertation

Abstract

Human action recognition from video requires representations that capture temporal dynamics while remaining computationally manageable. This thesis proposes a com- pact low-rank video representation framework for human action recognition based on temporal windowing and singular value decomposition (SVD). The central idea is to transform a sampled video sequence into a small set of informative component images that summarize dominant local temporal structure. Rather than processing the full video volume directly, the proposed framework applies uniform frame sampling, grayscale- difference construction, temporal windowing, matrix formulation, truncated SVD, and component-image generation before classification. The resulting compact representation is evaluated under two classification settings. In the first setting, the generated component images are used directly as compact descrip- tors and classified using logistic regression (LR) after standardization and principal component analysis (PCA). In the second setting, the component images are passed through a pretrained ResNet-50 backbone to extract deep visual features, which are then aggregated and classified using LR. The framework is evaluated primarily on the KTH benchmark using subject-aware Leave-One-Group-Out evaluation, with ad- ditional cross-dataset validation on UCF11 using a group-aware protocol. The experimental results show that the proposed representation is informative as a stan- dalone descriptor and becomes substantially more effective when combined with a pre- trained visual backbone. Across the ablation studies, the number of sampled frames emerges as the most influential design parameter, while the number of temporal win- dows and retained singular components further shape the trade-off between compact- ness and recognition performance. The strongest direct compact result is obtained with f64_w3_k3, while the strongest ResNet-based result is obtained with f64_w3_k4. These findings indicate that compact window-based low-rank summaries can preserve meaningful temporal and structural information for action recognition while providing a controllable alternative to heavier spatio-temporal pipelines.
Date of Award2026
Original languageAmerican English
Awarding Institution
  • HBKU College of Science and Engineering

Keywords

  • Compact video representation
  • Human action recognition
  • Low-rank video representation
  • ResNet-50 feature extraction
  • Singular value decomposition (SVD)
  • Temporal windowing

Cite this

'