Life sciences · Preprint
arXiv · September 8, 2026
Early or partial results. Treat as a signal, not a conclusion.
This is a descriptive benchmarking study comparing LLM inference performance across edge and near-edge hardware (Jetson Orin, CPU-only, and GPU-enabled server) using a fixed question-answering workload. The work characterizes trade-offs between latency, energy, accuracy, and model footprint but does not test a hypothesis or compare interventions; findings suggest that compute-side metrics alone may not optimally predict deployment outcomes for interactive services.
Controlled measurement study. Open-weight LLM models and their quantization variants. Intervention: LLM deployment across edge (Jetson Orin) and near-edge (server) hardware in CPU-only and GPU-enabled modes. Compared with: GPT-4o (cloud-hosted reference); cross-comparison of hardware configurations and quantization variants.
GPU-enabled server execution provides the lowest compute-side latency Jetson Orin shows lower measured energy consistent with lower platform power CPU-only execution is consistently dominated in latency and shows higher measured energy
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
A controlled measurement study of LLM inference across hardware platforms with multiple metrics, but no comparative intervention, randomization, or clinical outcome; findings are descriptive and exploratory rather than hypothesis-testing.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Large language models (LLMs) are increasingly used as backends for intelligent web services, but serving them across the edge continuum requires balancing quality, latency, model footprint, and energy. This paper presents a controlled measurement study of self-hosted LLM inference across edge and near-edge deployment nodes: an NVIDIA Jetson AGX Orin and a near-edge server with CPU-only and GPU-enabled inference modes. We evaluate multiple open-weight LLMs and quantization variants using a fixed question-answering workload, and compare them against GPT-4o as a cloud-hosted accuracy and latency reference. Our benchmarking pipeline reports accuracy, model footprint, per-token decoding latency, prefill latency, and overall execution energy. The results show that GPU-enabled server execution provides the lowest compute-side latency, while Jetson Orin shows lower measured energy, consistent with its lower platform power under our setup. CPU-only execution is consistently dominated in latency for our workload and shows higher measured energy. We also show that parameter count and downloaded weight-file size alone do not reliably predict observed accuracy or latency. Finally, using Pareto-frontier analysis, we study how deployment decisions may change under possible streamed-token delivery overheads, highlighting that compute-side inference metrics alone can lead to suboptimal placement for latency-sensitive interactive web services.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.