TY - JOUR
T1 - A naturalistic, non-invasive method for capturing biometric data during autism evaluations
AU - Kamal, Khaleel
AU - Hatvani, Janka
AU - Pethő, Máté
AU - Sárkány, András
AU - Hamvas, Imola
AU - Brune, Camille
AU - Wainer, Allison L.
AU - Dillon, Emily
AU - Berry Kravis, Elizabeth
AU - Ocampo, Edith Vanessa
AU - Arnold, Zachary Enos
AU - Ghazal, Iman
AU - Al-Faraj, Fatema
AU - Csákvári, Máté
AU - Katona-Pucsek, Kristóf
AU - Kun, Péter
AU - Oláh, Dóra
AU - Hernáth, Ferenc
AU - Schulc, Attila
AU - Mezősi, Anikó
AU - Latorre, Alejandro
AU - Al-Shaban, Fouad
AU - Soorya, Latha Valluripalli
AU - Tősér, Zoltán
N1 - Publisher Copyright:
Copyright © 2026 Kamal, Hatvani, Pethő, Sárkány, Hamvas, Brune, Wainer, Dillon, Berry Kravis, Ocampo, Arnold, Ghazal, Al-Faraj, Csákvári, Katona-Pucsek, Kun, Oláh, Hernáth, Schulc, Mezősi, Latorre, Al-Shaban, Soorya and Tősér.
PY - 2026/6/11
Y1 - 2026/6/11
N2 - Introduction – This study evaluated a machine learning tool designed to non-intrusively quantify and analyze biometric data of gaze, facial expressions, and paralinguistic social communication features during standardized autism observational assessments. The primary aim was to assess the diagnostic accuracy of this multimodal tool in capturing key social communication features of autism in a diverse neurodevelopmental disabilities cohort and neurotypical (NT) cohort, ages 2-12. Methods – The study enrolled 546 participants across four sites in the USA (n=246) and Qatar (n=300). Of these, 458 (83.6%) met quality indicators for both video and audio recordings and comprised the analysis set. Primary outcome measures were diagnostic accuracy, sensitivity and specificity relative to reference diagnoses. Random Forest classifiers were trained using a developmentally-adaptive approach: separate models for three developmental groups (few-to-no words, phrase speech, fluent speech) using 97 biometric features from video, audio and gaze data. Performance was assessed via leave-one-out cross-validation on the inner set (n=338) and validated on an independent hold-out test set (n=120). Results – Classification between ASD (idiopathic and syndromic) and non-ASD (clinical and NT) participants achieved 77.8% sensitivity and specificity in the inner set. When distinguishing ASD from NT participants alone, the sensitivity and specificity both increased to 82.0%. In the hold-out test set, the model demonstrated 62.3% sensitivity and 81.4% specificity for ASD versus non-ASD, and 72.1% sensitivity and 88.6% specificity for ASD versus NT. Performance varied by demographics in the case of sex, with males showing higher sensitivity (79% to 75%) and females showing higher specificity (84% to 70%). Discussion – This study demonstrates the feasibility of using semi-automated multimodal computational analysis to quantify multi-modal autism social communication behaviors and distinguish ASD from NT in clinically and ethnically diverse samples. Known difficulties with differential diagnosis in non-ASD neurodevelopmental conditions with autism-like features remain a limitation requiring further development. Data suggest promise for such tools to support task-sharing models within existing clinical approaches.
AB - Introduction – This study evaluated a machine learning tool designed to non-intrusively quantify and analyze biometric data of gaze, facial expressions, and paralinguistic social communication features during standardized autism observational assessments. The primary aim was to assess the diagnostic accuracy of this multimodal tool in capturing key social communication features of autism in a diverse neurodevelopmental disabilities cohort and neurotypical (NT) cohort, ages 2-12. Methods – The study enrolled 546 participants across four sites in the USA (n=246) and Qatar (n=300). Of these, 458 (83.6%) met quality indicators for both video and audio recordings and comprised the analysis set. Primary outcome measures were diagnostic accuracy, sensitivity and specificity relative to reference diagnoses. Random Forest classifiers were trained using a developmentally-adaptive approach: separate models for three developmental groups (few-to-no words, phrase speech, fluent speech) using 97 biometric features from video, audio and gaze data. Performance was assessed via leave-one-out cross-validation on the inner set (n=338) and validated on an independent hold-out test set (n=120). Results – Classification between ASD (idiopathic and syndromic) and non-ASD (clinical and NT) participants achieved 77.8% sensitivity and specificity in the inner set. When distinguishing ASD from NT participants alone, the sensitivity and specificity both increased to 82.0%. In the hold-out test set, the model demonstrated 62.3% sensitivity and 81.4% specificity for ASD versus non-ASD, and 72.1% sensitivity and 88.6% specificity for ASD versus NT. Performance varied by demographics in the case of sex, with males showing higher sensitivity (79% to 75%) and females showing higher specificity (84% to 70%). Discussion – This study demonstrates the feasibility of using semi-automated multimodal computational analysis to quantify multi-modal autism social communication behaviors and distinguish ASD from NT in clinically and ethnically diverse samples. Known difficulties with differential diagnosis in non-ASD neurodevelopmental conditions with autism-like features remain a limitation requiring further development. Data suggest promise for such tools to support task-sharing models within existing clinical approaches.
KW - Asd
KW - Autism spectrum disorder
KW - Behavioral analysis
KW - Biometric data
KW - Diagnostic support
KW - Machine learning
KW - Pediatrics
UR - https://www.scopus.com/pages/publications/105044396453
U2 - 10.3389/fpsyt.2026.1819384
DO - 10.3389/fpsyt.2026.1819384
M3 - Article
AN - SCOPUS:105044396453
SN - 1664-0640
VL - 17
JO - Frontiers in Psychiatry
JF - Frontiers in Psychiatry
M1 - 1819384
ER -