Human action recognition from video requires representations that capture temporal
dynamics while remaining computationally manageable. This thesis proposes a com-
pact low-rank video representation framework for human action recognition based on
temporal windowing and singular value decomposition (SVD). The central idea is to
transform a sampled video sequence into a small set of informative component images
that summarize dominant local temporal structure. Rather than processing the full video
volume directly, the proposed framework applies uniform frame sampling, grayscale-
difference construction, temporal windowing, matrix formulation, truncated SVD, and
component-image generation before classification.
The resulting compact representation is evaluated under two classification settings. In
the first setting, the generated component images are used directly as compact descrip-
tors and classified using logistic regression (LR) after standardization and principal
component analysis (PCA). In the second setting, the component images are passed
through a pretrained ResNet-50 backbone to extract deep visual features, which are
then aggregated and classified using LR. The framework is evaluated primarily on
the KTH benchmark using subject-aware Leave-One-Group-Out evaluation, with ad-
ditional cross-dataset validation on UCF11 using a group-aware protocol.
The experimental results show that the proposed representation is informative as a stan-
dalone descriptor and becomes substantially more effective when combined with a pre-
trained visual backbone. Across the ablation studies, the number of sampled frames
emerges as the most influential design parameter, while the number of temporal win-
dows and retained singular components further shape the trade-off between compact-
ness and recognition performance. The strongest direct compact result is obtained with
f64_w3_k3, while the strongest ResNet-based result is obtained with f64_w3_k4.
These findings indicate that compact window-based low-rank summaries can preserve
meaningful temporal and structural information for action recognition while providing
a controllable alternative to heavier spatio-temporal pipelines.
| Date of Award | 2026 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - HBKU College of Science and Engineering
|
|---|
- Compact video representation
- Human action recognition
- Low-rank video representation
- ResNet-50 feature extraction
- Singular value decomposition (SVD)
- Temporal windowing
A LOW-RANK VIDEO REPRESENTATION FRAMEWORK FOR HUMAN ACTION RECOGNITION
Al-Amir, L. (Author). 2026
Student thesis: Master's Dissertation