Characterizing Prefill and Decode Regimes for Edge LLM Inference
2026 49th MIPRO ICT and Electronics Convention (MIPRO). IEEE 2026 S. 123 - 128
Erscheinungsjahr: 2026
Publikationstyp: Zeitschriftenaufsatz
Sprache: Englisch
Doi/URN: 10.1109/mipro70003.2026.11591742
| Geprüft: | Bibliothek |
Inhaltszusammenfassung
Large language models (LLMs) are increasingly deployed on edge devices to reduce latency, preserve privacy, and enable offline operation. However, edge platforms impose strict constraints on memory capacity, power consumption, and thermal dissipation, which shape inference behavior. Modern LLM inference consists of distinct execution regimes—prompt evaluation and autoregressive decoding—with different scaling and stability characteristics, yet are often assessed using aggregate performance me...Large language models (LLMs) are increasingly deployed on edge devices to reduce latency, preserve privacy, and enable offline operation. However, edge platforms impose strict constraints on memory capacity, power consumption, and thermal dissipation, which shape inference behavior. Modern LLM inference consists of distinct execution regimes—prompt evaluation and autoregressive decoding—with different scaling and stability characteristics, yet are often assessed using aggregate performance metrics.In this work, we experimentally characterize LLM inference on representative edge hardware by isolating prefill and decode behavior under input length scaling and sustained execution. We evaluate a CPU-only Raspberry Pi 5 and a GPU-enabled NVIDIA Jetson Orin Nano using a common inference runtime and quantized LLaMA models. Our results show that increasing input length degrades both prefill and decode performance, with decode throughput decreasing as context length grows and thermal effects amplifying this degradation on CPU-only platforms. Under sustained workloads, the Jetson Orin Nano maintains stable performance and favorable energy efficiency, while the Raspberry Pi 5 exhibits throughput and efficiency degradation due to passive cooling and thermal throttling.These results underscore the importance of regime-aware evaluation and thermal considerations for practical LLM deployment on edge devices.» weiterlesen» einklappen
Schlüsselwörter
Autoren
Klassifikation
DFG Fachgebiet:
4.43-04 - Künstliche Intelligenz und Maschinelle Lernverfahren
DDC Sachgruppe:
Informatik
Verknüpfte Personen
- Stefan Naumann
- Professor
(Umweltplanung/Umwelttechnik)
- Florian Mohr
- Professor
(Umweltplanung/Umwelttechnik)