TY - GEN
T1 - Think Fast, Infer Smart
T2 - 36th IEEE International Symposium on Personal, Indoor and Mobile Radio Communications, PIMRC 2025
AU - Albaseer, Abdullatif
AU - Bentafat, Elmahdi
AU - Hamood, Moqbel
AU - Abdallah, Mohamed
AU - Al-Fuqaha, Ala
AU - Hamdi, Mounir
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025/9/4
Y1 - 2025/9/4
N2 - Deploying large language models (LLMs) at the wireless edge is a promising solution to meet the low-latency, high-computation demands of next-generation AI applications. Although the existing literature has introduced approaches to enable distributed LLM inference, these methods largely overlook the distinct computational and communication characteristics of the two-phase LLM inference process - the pre-fill and decode phases. This oversight leads to suboptimal performance and limits scalability in real-world deployments. To address these issues, we propose a novel collaborative inference framework that strategically minimizes inference latency by optimally distributing computational loads across edge devices, the edge server, and the cloud. Our approach introduces a hybrid framework that combines head-wise parallel processing with layer-wise partitioning of LLM models, supported by a dual-phase optimization strategy. In the pre-fill phase, we optimize assigning attention heads to selected edge devices for parallel computation and efficient resource use. We then optimize for minimal latency by selecting participants, determining head assignments per device, and allocating bandwidth while meeting all constraints. In the de-code phase, our framework adaptively decides whether to execute computations locally on the edge server, offload them to the cloud, or redistribute tasks among edge devices, optimizing this decision based on the remaining latency budget and the sequential nature of the decode phase. The simulation results demonstrate that the proposed framework significantly outperforms the baseline methods, achieving a 56% reduction in inference latency, 40% improvement in bandwidth efficiency and 35% improvement in resource utilization.
AB - Deploying large language models (LLMs) at the wireless edge is a promising solution to meet the low-latency, high-computation demands of next-generation AI applications. Although the existing literature has introduced approaches to enable distributed LLM inference, these methods largely overlook the distinct computational and communication characteristics of the two-phase LLM inference process - the pre-fill and decode phases. This oversight leads to suboptimal performance and limits scalability in real-world deployments. To address these issues, we propose a novel collaborative inference framework that strategically minimizes inference latency by optimally distributing computational loads across edge devices, the edge server, and the cloud. Our approach introduces a hybrid framework that combines head-wise parallel processing with layer-wise partitioning of LLM models, supported by a dual-phase optimization strategy. In the pre-fill phase, we optimize assigning attention heads to selected edge devices for parallel computation and efficient resource use. We then optimize for minimal latency by selecting participants, determining head assignments per device, and allocating bandwidth while meeting all constraints. In the de-code phase, our framework adaptively decides whether to execute computations locally on the edge server, offload them to the cloud, or redistribute tasks among edge devices, optimizing this decision based on the remaining latency budget and the sequential nature of the decode phase. The simulation results demonstrate that the proposed framework significantly outperforms the baseline methods, achieving a 56% reduction in inference latency, 40% improvement in bandwidth efficiency and 35% improvement in resource utilization.
KW - and bandwidth optimization
KW - dynamic head-wise layer-wise partitioning
KW - Edge computing
KW - Large Language Models (LLMs)
KW - latency
UR - https://www.scopus.com/pages/publications/105030546202
U2 - 10.1109/PIMRC62392.2025.11274878
DO - 10.1109/PIMRC62392.2025.11274878
M3 - Conference contribution
AN - SCOPUS:105030546202
T3 - IEEE International Symposium on Personal, Indoor and Mobile Radio Communications, PIMRC
BT - 2025 IEEE 36th International Symposium on Personal, Indoor and Mobile Radio Communications, PIMRC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 1 September 2025 through 4 September 2025
ER -