TY - GEN
T1 - Learning to Control Dynamical Agents via Spiking Neural Networks and Metropolis-Hastings Sampling
AU - Safa, Ali
AU - Mohsen, Farida
AU - Al-Zawqari, Ali
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Spiking Neural Networks (SNNs) offer biologically inspired, energy-efficient alternatives to traditional Deep Neural Networks (DNNs) for real-Time control systems. However, their training presents several challenges, particularly for reinforcement learning (RL) tasks, due to the non-differentiable nature of spike-based communication. In this work, we introduce what is, to our knowledge, the first framework that employs Metropolis-Hastings (MH) sampling, a Bayesian inference technique, to train SNNs for dynamical agent control in RL environments without relying on gradient-based methods. Our approach iteratively proposes and probabilistically accepts network parameter updates based on accumulated reward signals, effectively circum-venting the limitations of backpropagation while enabling direct optimization on neuromorphic platforms. We evaluated this framework on two standard control benchmarks: AcroBot and CartPole. The results demonstrate that our MH-based approach outperforms conventional Deep Q-Learning (DQL) baselines and prior SNN-based RL approaches in terms of maximizing the accumulated reward while minimizing network resources and training episodes.
AB - Spiking Neural Networks (SNNs) offer biologically inspired, energy-efficient alternatives to traditional Deep Neural Networks (DNNs) for real-Time control systems. However, their training presents several challenges, particularly for reinforcement learning (RL) tasks, due to the non-differentiable nature of spike-based communication. In this work, we introduce what is, to our knowledge, the first framework that employs Metropolis-Hastings (MH) sampling, a Bayesian inference technique, to train SNNs for dynamical agent control in RL environments without relying on gradient-based methods. Our approach iteratively proposes and probabilistically accepts network parameter updates based on accumulated reward signals, effectively circum-venting the limitations of backpropagation while enabling direct optimization on neuromorphic platforms. We evaluated this framework on two standard control benchmarks: AcroBot and CartPole. The results demonstrate that our MH-based approach outperforms conventional Deep Q-Learning (DQL) baselines and prior SNN-based RL approaches in terms of maximizing the accumulated reward while minimizing network resources and training episodes.
KW - Control
KW - Dynamical Agent
KW - Metropolis-Hastings sampling
KW - Reinforcement Learning
KW - Spiking Neural Networks
UR - https://www.scopus.com/pages/publications/105033970309
U2 - 10.1109/MIND67540.2025.11351545
DO - 10.1109/MIND67540.2025.11351545
M3 - Conference contribution
AN - SCOPUS:105033970309
T3 - Proceedings of the 2025 International Conference on Machine Intelligence and Nature-Inspired Computing, MIND 2025
SP - 1
EP - 6
BT - Proceedings of the 2025 International Conference on Machine Intelligence and Nature-Inspired Computing, MIND 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 International Conference on Machine Intelligence and Nature-Inspired Computing, MIND 2025
Y2 - 31 October 2025 through 2 November 2025
ER -