Skip to main navigation Skip to search Skip to main content

Think Fast, Infer Smart: A Hybrid Distributed LLMs Inference at the Wireless Edge

  • Hamad bin Khalifa University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Deploying large language models (LLMs) at the wireless edge is a promising solution to meet the low-latency, high-computation demands of next-generation AI applications. Although the existing literature has introduced approaches to enable distributed LLM inference, these methods largely overlook the distinct computational and communication characteristics of the two-phase LLM inference process - the pre-fill and decode phases. This oversight leads to suboptimal performance and limits scalability in real-world deployments. To address these issues, we propose a novel collaborative inference framework that strategically minimizes inference latency by optimally distributing computational loads across edge devices, the edge server, and the cloud. Our approach introduces a hybrid framework that combines head-wise parallel processing with layer-wise partitioning of LLM models, supported by a dual-phase optimization strategy. In the pre-fill phase, we optimize assigning attention heads to selected edge devices for parallel computation and efficient resource use. We then optimize for minimal latency by selecting participants, determining head assignments per device, and allocating bandwidth while meeting all constraints. In the de-code phase, our framework adaptively decides whether to execute computations locally on the edge server, offload them to the cloud, or redistribute tasks among edge devices, optimizing this decision based on the remaining latency budget and the sequential nature of the decode phase. The simulation results demonstrate that the proposed framework significantly outperforms the baseline methods, achieving a 56% reduction in inference latency, 40% improvement in bandwidth efficiency and 35% improvement in resource utilization.

Original languageEnglish
Title of host publication2025 IEEE 36th International Symposium on Personal, Indoor and Mobile Radio Communications, PIMRC 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798350363234
DOIs
Publication statusPublished - 4 Sept 2025
Event36th IEEE International Symposium on Personal, Indoor and Mobile Radio Communications, PIMRC 2025 - Istanbul, Turkey
Duration: 1 Sept 20254 Sept 2025

Publication series

NameIEEE International Symposium on Personal, Indoor and Mobile Radio Communications, PIMRC
ISSN (Print)2166-9570
ISSN (Electronic)2166-9589

Conference

Conference36th IEEE International Symposium on Personal, Indoor and Mobile Radio Communications, PIMRC 2025
Country/TerritoryTurkey
CityIstanbul
Period1/09/254/09/25

Keywords

  • and bandwidth optimization
  • dynamic head-wise layer-wise partitioning
  • Edge computing
  • Large Language Models (LLMs)
  • latency

Fingerprint

Dive into the research topics of 'Think Fast, Infer Smart: A Hybrid Distributed LLMs Inference at the Wireless Edge'. Together they form a unique fingerprint.

Cite this