Starten Sie Ihre Suche...


Wir weisen darauf hin, dass wir technisch notwendige Cookies verwenden. Weitere Informationen

Characterizing Prefill and Decode Regimes for Edge LLM Inference

2026 49th MIPRO ICT and Electronics Convention (MIPRO). IEEE 2026 S. 123 - 128

Erscheinungsjahr: 2026

Publikationstyp: Zeitschriftenaufsatz

Sprache: Englisch

Doi/URN: 10.1109/mipro70003.2026.11591742

Volltext über DOI/URN

Geprüft:Bibliothek

Inhaltszusammenfassung


Large language models (LLMs) are increasingly deployed on edge devices to reduce latency, preserve privacy, and enable offline operation. However, edge platforms impose strict constraints on memory capacity, power consumption, and thermal dissipation, which shape inference behavior. Modern LLM inference consists of distinct execution regimes—prompt evaluation and autoregressive decoding—with different scaling and stability characteristics, yet are often assessed using aggregate performance me...Large language models (LLMs) are increasingly deployed on edge devices to reduce latency, preserve privacy, and enable offline operation. However, edge platforms impose strict constraints on memory capacity, power consumption, and thermal dissipation, which shape inference behavior. Modern LLM inference consists of distinct execution regimes—prompt evaluation and autoregressive decoding—with different scaling and stability characteristics, yet are often assessed using aggregate performance metrics.In this work, we experimentally characterize LLM inference on representative edge hardware by isolating prefill and decode behavior under input length scaling and sustained execution. We evaluate a CPU-only Raspberry Pi 5 and a GPU-enabled NVIDIA Jetson Orin Nano using a common inference runtime and quantized LLaMA models. Our results show that increasing input length degrades both prefill and decode performance, with decode throughput decreasing as context length grows and thermal effects amplifying this degradation on CPU-only platforms. Under sustained workloads, the Jetson Orin Nano maintains stable performance and favorable energy efficiency, while the Raspberry Pi 5 exhibits throughput and efficiency degradation due to passive cooling and thermal throttling.These results underscore the importance of regime-aware evaluation and thermal considerations for practical LLM deployment on edge devices.» weiterlesen» einklappen

Schlüsselwörter


  • Edge AI
  • Large Language Models
  • Embedded Systems
  • LLM Inference
  • Performance Evaluation

Autoren


Scheffler, Jonas (Autor)
Fazlic, Lejla Begic (Autor)

Klassifikation


DFG Fachgebiet:
4.43-04 - Künstliche Intelligenz und Maschinelle Lernverfahren

DDC Sachgruppe:
Informatik

Verknüpfte Personen



Beteiligte Einrichtungen