TY - GEN
T1 - Real-Time Human Interaction Intent Detection for Resource-Constrained Social Robots Using RGB-Only Behavioral and Emotional Cues
AU - Mohsen, Farida
AU - Safa, Ali
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026/3/12
Y1 - 2026/3/12
N2 - Service robots with expressive heads are increasingly being deployed in public spaces, creating a need for the real-time detection of human intent to interact, in order to provide responsive services back to the interacting human users. In this work, we present a practical and modular deep learning-based system for detecting human interaction intent, which runs in real time on compute- and resource-limited robot platforms without requiring high-cost, bulky, and power-hungry GPUs. Our approach combines human pose estimation and facial emotion analysis for both frame-level and sequence-level detection of interaction-intent behaviors. Crucially, we demonstrate robust model generalization across different cameras and environments by implementing our system in hardware (Raspberry Pi 5) and using it in the field. Remarkably, we show that even though our dataset has been collected in a controlled setting using a standard USB webcam, our trained model generalizes successfully without finetuning on a robot head platform which makes use of a different camera sensor and deployed in different uncontrolled environments not captured by the training data. In contrast to most prior works that report model precision in offline settings only, our real-world system deployment and assessment in the wild reports a high balanced detection accuracy of 91%, clearly demonstrating the usefulness of our approach in practice.
AB - Service robots with expressive heads are increasingly being deployed in public spaces, creating a need for the real-time detection of human intent to interact, in order to provide responsive services back to the interacting human users. In this work, we present a practical and modular deep learning-based system for detecting human interaction intent, which runs in real time on compute- and resource-limited robot platforms without requiring high-cost, bulky, and power-hungry GPUs. Our approach combines human pose estimation and facial emotion analysis for both frame-level and sequence-level detection of interaction-intent behaviors. Crucially, we demonstrate robust model generalization across different cameras and environments by implementing our system in hardware (Raspberry Pi 5) and using it in the field. Remarkably, we show that even though our dataset has been collected in a controlled setting using a standard USB webcam, our trained model generalizes successfully without finetuning on a robot head platform which makes use of a different camera sensor and deployed in different uncontrolled environments not captured by the training data. In contrast to most prior works that report model precision in offline settings only, our real-world system deployment and assessment in the wild reports a high balanced detection accuracy of 91%, clearly demonstrating the usefulness of our approach in practice.
KW - human-robot interaction (HRI)
KW - interaction intent detection
KW - real-time inference
KW - resource-constrained deployment
KW - Social service robots
UR - https://www.scopus.com/pages/publications/105043995992
U2 - 10.1109/ABC68169.2026.11567041
DO - 10.1109/ABC68169.2026.11567041
M3 - Conference contribution
AN - SCOPUS:105043995992
T3 - 8th International Conference on Activity and Behavior Computing, ABC 2026
BT - 8th International Conference on Activity and Behavior Computing, ABC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 8th International Conference on Activity and Behavior Computing, ABC 2026
Y2 - 9 March 2026 through 12 March 2026
ER -