Abstract
Deploying deepfake detection in real-world settings is challenging because the most valuable training data is distributed across devices and organizations (e.g., institutions, platforms, and users) and cannot be centralized freely due to privacy and governance restrictions. Federated learning (FL) enables learning from such distributed video data without exposing raw media. Yet, current reconstruction-based federated deepfake detectors usually aggregate frame evidence using uniform temporal pooling, which assumes equal contribution from all frames. Since deepfake artifacts are often sparse and temporally localized, uniform pooling can weaken the supervision signal and reduce detection reliability. Therefore, in this work, we propose Residual-Guided Temporal Pooling (RGTP), a reconstruction-based federated deepfake detection framework that leverages frame-level reconstruction residuals to drive adaptive temporal aggregation during local training. RGTP assigns higher weights to frames with larger residual magnitudes, producing more informative video-level logits while preserving the standard FL communication pattern and the data-locality properties of FL. To improve stability under noisy or low-quality frames, we further introduce a temporal consistency regularizer that encourages coherent frame-level predictions across time. Experiments on Celeb-DF and FaceForensics++ show consistent gains over uniform pooling across AUROC, AUPRC, and F1 score. Because RGTP operates only inside each client’s local temporal aggregation step, it does not add client-server communication; the additional local computation is limited to residual normalization, softmax weighting, and the temporal consistency term over already computed frame-level outputs.
| Original language | English |
|---|---|
| Pages (from-to) | 85109-85120 |
| Number of pages | 12 |
| Journal | IEEE Access |
| Volume | 14 |
| DOIs | |
| Publication status | Published - 2026 |
Keywords
- Conferences
- Deepfake detection
- Deepfakes
- Faces
- Federated learning
- Forensics
- Forgery
- Modeling
- Privacy-preserving learning
- Reconstruction-based detection
- Signal detection
- Temporal pooling
- Training
- Video forensics
- Videos
Fingerprint
Dive into the research topics of 'RGTP: A Residual-Guided Temporal Pooling Framework for Federated Deepfake Video Detection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver