Software-Defined
Stochastic Inference Engine

Entropy-governed dynamic precision routing and speculative verification for energy-constrained large language model inference on modern datacenter and edge accelerators.

This is the broader, ongoing research project. Its most thoroughly measured component — fixed-K5 speculative decoding — has its own repository, now with a like-for-like, energy-measured comparison against EAGLE-3 inside vLLM and SGLang (below).

System Architecture

The Hardware-Aware Adaptive Runtime

Decoupling wall-clock execution time from deterministic silicon precision without paying the software emulation tax.

1

Shannon Entropy Routing

Calculates instantaneous logit uncertainty in real time (Ht = −∑ pi log2 pi). The concept: route routine syntax and predictable tokens to sub-byte execution, reserving full FP16 precision for complex reasoning transitions. Now wired to real computation for the first time — and found to have a sharp fidelity limit at more than one simultaneously-gated layer (see Broader Research Project below).

Ht < θlow → Shift to INT4 High Gear
2

Schmitt-Trigger Clutch

Prevents pipeline thrashing and "gear hunting" across sentence boundaries via a dual-threshold deadband (θlow, θhigh, h) and smoothed moving entropy window (W = 3).

Hysteresis deadband prevents stalls
3

Speculative Verification

Pairs a lightweight draft model (1B) with a target verifier (8B, the only size tested so far) in a unified VRAM pool. Draft acceptance ranges from 42.5% to 85.4% depending on how predictable the text is (cache-free harness, N=10 trials/prompt), validating several tokens in a single target pass.

1-Pass parallel multi-token validation

Datacenter Scale

Why There Is No Savings Calculator (Yet)

An earlier version of this page projected datacenter energy and cost savings from our measurements. It has been withdrawn, because every result so far was measured with one request at a time (batch size 1) on a single consumer GPU.

Datacenters serve many requests at once, and the benefit of speculative decoding is known to shrink as batches grow: the GPU capacity it relies on is increasingly used by the other requests. Scaling a batch-1 measurement up to a whole cluster would therefore overstate the savings.

What we have measured: at batch size 1, fixed-K5 speculative decoding cuts GPU energy per token by 53–55% inside vLLM and SGLang. The next step is measuring how that changes at realistic batch sizes. A projection will return here only once it rests on those measurements.

LIKE-FOR-LIKE BENCHMARK · v3.0.1

Fixed-K=5 Speculative Decoding vs. EAGLE-3

A small draft model (Llama-3.2-1B) proposes five tokens at a time and the target model (Llama-3.1-8B) verifies them in one pass. This is the standard draft-model method, not a new algorithm; the contribution is careful measurement. Tested inside vLLM and SGLang against EAGLE-3, EAGLE-1 and n-gram drafting, each at its best configuration, on 27 prompts, with GPU energy read from the hardware energy counter.

Result: fixed-K5 roughly halves energy per token and raises throughput by 60–79%. A well-configured EAGLE-3 is faster still, by 8–14%, with energy per token within about 10% either way. Fixed-K5's advantage is that it needs no specially trained draft head, so it also works for fine-tuned, custom and newly released models. Batch size 1 only; see the repository for scope and limitations.

−53–55%
Energy per token
+60–79%
Throughput
8–14%
Slower than best EAGLE-3
27×2
Prompts × engines

Broader Research Project

Component-level results from the ongoing research repository. Some components are validated; the entropy-gated adaptive mechanisms below have real, characterized limitations — see the research repo README for the full account.

Energy Expenditure
−53% to −55%

Per-token GPU energy reduction from fixed-K5 draft-model speculative decoding inside vLLM and SGLang (batch size 1, 27 prompts). In the separate cache-free harness: −32.7% to −60.8%, depending on how predictable the text is. See the fixed-K5 repository for methodology and raw telemetry.

Memory Bus Traffic
-71.9%

Memory bandwidth cut per projection layer via custom Triton INT4 GEMM kernel.

Resolution Gear Fidelity Limit
1 layer safe, 2+ collapses

Entropy-gated INT4/FP16 precision switching is now wired to real computation (not just logged): output matches an FP16 baseline closely at a single gated decoder layer, but fidelity collapses sharply (7–10% token match) at two or more — a real, characterized limitation, not a bug. Full data in the research repo README.

Kernel Latency
76.5 µs

Isolated INT4 GEMM latency, real branching kernel test (M=1). Only ~4% faster than FP16 at this batch size — bandwidth savings do not yet translate to proportional latency gains; see the research repo for details.

Academic Citation

Cite This Work

Permanent DOI minted via CERN / Zenodo. It always resolves to the latest version of the fixed-K5 release.

@software{jacklin2026fixedk5,
  author       = {Zanno Jacklin},
  title        = {{Fixed-K Speculative Decoding vs. EAGLE-3: A Like-for-Like,
                   Energy-Measured Comparison in vLLM and SGLang}},
  month        = sep,
  year         = 2026,
  publisher    = {Zenodo},
  version      = {v3.0.1},
  doi          = {10.5281/zenodo.22210487},
  url          = {https://doi.org/10.5281/zenodo.22210487}
}