TY - GEN
T1 - Dynamic input anomaly detection in interactive multimedia services
AU - Shatnawi, Mohammed
AU - Hefeeda, Mohamed
N1 - Publisher Copyright:
© 2018 Association for Computing Machinery.
PY - 2018/6/12
Y1 - 2018/6/12
N2 - Multimedia services like Skype, WhatsApp, and Google Hangouts have strict Service Level Agreements (SLAs). These services attempt to address the root causes of SLA violations through techniques such as detecting anomalies in the inputs of the services. The key problem with current anomaly detection and handling techniques is that they can't adapt to service changes in real-time. In current techniques, historic data from prior runs of the service are used to identify anomalies in the service inputs like number of concurrent users, and system states like CPU utilization. These techniques do not evaluate the current impact of anomalies on the service. Thus, they may raise alerts and take corrective measures even if the detected anomalies do not cause SLA violations. Alerts are expensive to handle from a system and engineering support perspectives, and should be raised only if necessary. We propose a dynamic approach for handling service input and system state anomalies in multimedia services in real-time, by evaluating the impact of anomalies, independently and associatively, on the service outputs. Our proposed approach alerts and takes corrective measures like capacity allocations if the detected anomalies result in SLA violations.We implement our approach in a large-scale operational multimedia service, and show that it increases anomaly detection accuracy by 31%, reduces anomaly alerting false positives by 71%, false negatives by 69%, and enhances media sharing quality by 14%.
AB - Multimedia services like Skype, WhatsApp, and Google Hangouts have strict Service Level Agreements (SLAs). These services attempt to address the root causes of SLA violations through techniques such as detecting anomalies in the inputs of the services. The key problem with current anomaly detection and handling techniques is that they can't adapt to service changes in real-time. In current techniques, historic data from prior runs of the service are used to identify anomalies in the service inputs like number of concurrent users, and system states like CPU utilization. These techniques do not evaluate the current impact of anomalies on the service. Thus, they may raise alerts and take corrective measures even if the detected anomalies do not cause SLA violations. Alerts are expensive to handle from a system and engineering support perspectives, and should be raised only if necessary. We propose a dynamic approach for handling service input and system state anomalies in multimedia services in real-time, by evaluating the impact of anomalies, independently and associatively, on the service outputs. Our proposed approach alerts and takes corrective measures like capacity allocations if the detected anomalies result in SLA violations.We implement our approach in a large-scale operational multimedia service, and show that it increases anomaly detection accuracy by 31%, reduces anomaly alerting false positives by 71%, false negatives by 69%, and enhances media sharing quality by 14%.
KW - Multimedia communication services
KW - Multimedia service anomaly detection
KW - Multimedia service reliability
UR - https://www.scopus.com/pages/publications/85050649728
U2 - 10.1145/3204949.3204954
DO - 10.1145/3204949.3204954
M3 - Conference contribution
AN - SCOPUS:85050649728
T3 - Proceedings of the 9th ACM Multimedia Systems Conference, MMSys 2018
SP - 113
EP - 122
BT - Proceedings of the 9th ACM Multimedia Systems Conference, MMSys 2018
PB - Association for Computing Machinery, Inc
T2 - 9th ACM Multimedia Systems Conference, MMSys 2018
Y2 - 12 June 2018 through 15 June 2018
ER -