12 min read
Suppose you're building an agent that works with private files.
You moved inference onto the laptop for privacy, split the work into an orchestrator plus a few workers, and turned on speculative decoding.
It still feels slow.
The model is probably fine but the problem is that the inference engines we use locally were built for one user making one request at a time, and a multi-agent workflow doesn't look like that.
EdgeAgent takes this problem apart carefully.
On Apple M4, there is a 1.29x speedup from UMA-aware execution alone, compared with batched speculative decoding. The full system reaches 1.77x under extreme tool-use latency.
Let's have a look at why the problem exists and what's the solution.
On Apple silicon, the CPU and GPU share one memory pool. MLX's unified memory docs put it plainly: "The CPU and GPU have direct access to the same memory pool."
That makes zero-copy sharing possible, it also means both processors compete for the same bus.

Decoding one token at a time is limited by memory bandwidth. If you need a refresher on why, see why memory bandwidth, not math, makes LLM inference slow.
Micro-benchmarks on the M4 show what happens when CPU and GPU run together:
| Matrix (K×N) | CPU GB/s | GPU GB/s | Co-run GB/s | Co-run vs GPU |
|---|---|---|---|---|
| 4096×4096 | 66.1 | 54.3 | 73.9 | 1.36x |
| 4096×14336 | 68.9 | 80.3 | 91.3 | 1.14x |
| 4096×128256 | 72.0 | 96.7 | 95.4 | 0.99x |
On the vocabulary-sized matrix, the GPU alone already uses about 80% of the 120 GB/s peak. Adding the CPU makes things slightly worse, so splitting decode naively across CPU and GPU gives you almost nothing.

Speculative decoding helps because it loads the target model's weights once for several candidate tokens.
How much it helps depends on how predictable the output is.
EdgeAgent's M4 experiment with DeepSeek-R1-Distill-Llama-8B, structured output peaks at 1.69x with draft length 16 and then gets worse as rejections pile up.
Reasoning tasks hit diminishing returns at much shorter draft lengths.

In a multi-agent system, your orchestrator writes open-ended reasoning, which is hard to draft for.
Your workers produce code and tool-call JSON, which are easy.
Current speculative decoding systems use one draft length for the whole batch, that wastes bandwidth on the orchestrator's rejected drafts and limits the workers.
For more on how different speculators compare, and the acceptance rate below which speculation costs more than it saves, see the EAGLE vs Medusa vs draft-model comparison.
Workers spend time waiting on scripts and APIs.
If a stalled agent still holds slots in the batch, those slots do nothing. Static batching assumes every sequence is always ready to run, and agent workloads break that assumption.

This is already visible in shipping tools. LM Studio's MLX engine fails with SpeculativeDecodingNotSupportedError when you combine a draft model with batched MLX models. "Batched" and "speculative" are still treated as features you can't use together, and multi-agent workloads need both.
EdgeAgent's CPU path depends on Arm's Scalable Matrix Extension (SME). The Hello SME paper reports that the M4 was the first chip to support SME.
$ clang -O2 -march=armv8-a+sme -o veclen veclen.c
$ ./veclen
Has SME? Yes
Streaming mode: svcntw() = 16, svcntsw() = 16If you're on M1 through M3, EdgeAgent's CPU kernels don't apply to your machine.
The MLX docs show the programming model. Arrays live in unified memory, and you pick the device per operation instead of moving data:
a = mx.random.normal((100,))
b = mx.random.normal((100,))
mx.add(a, b, stream=mx.cpu)
mx.add(a, b, stream=mx.gpu)When one stream depends on another, MLX inserts the synchronization for you:
c = mx.add(a, b, stream=mx.cpu)
d = mx.add(a, c, stream=mx.gpu)This automatic tracking works at the level of whole tensors, and EdgeAgent runs into its limits. When the CPU and GPU write separate column ranges of the same buffer, the graph can't tell the writes don't overlap. It treats them as a conflict and adds conservative synchronization, or it forces a concat that copies both halves. I cover the fix below.
The mlx-lm server supports speculative decoding through --draft-model and --num-draft-tokens. Measure with and without these flags before you assume they help. One mlx-lm issue reports a 35% throughput drop on a MoE target (Qwen3.5-397B-A17B, with Qwen3.5-9B as the draft) because 17B active parameters is too close to the draft model's size. It's the same lesson as EdgeAgent: speculation only pays off when the workload suits it.
EdgeAgent fixes partitioning when the model loads:
n = α·N columns and the GPU gets the rest. α is computed per layer from K/N so both devices finish at about the same time.Because each device writes its own columns, you never need a reduction across devices. Here's a NumPy sketch of the idea (mine, not the paper's):

import numpy as np
K, N = 4096, 14336 # an MLP up-projection shape from the paper's Table 1
alpha = 0.4 # ASSUMPTION: paper derives alpha per layer from K/N
n = int(alpha * N)
x = np.random.randn(1, K).astype(np.float32) # one decode token
W = np.random.randn(K, N).astype(np.float32)
W_cpu, W_gpu = W[:, :n], W[:, n:] # split on output columns
out = np.empty((1, N), dtype=np.float32) # one preallocated buffer
out[:, :n] = x @ W_cpu # "CPU" writes its columns
out[:, n:] = x @ W_gpu # "GPU" writes the rest
print(np.allclose(out, x @ W, rtol=1e-4, atol=1e-3)) # True: no reduction neededThe hard part is the graph.
EdgeAgent replaces concat with a Logical Dependency Barrier. It's a node that makes downstream operators wait for both producers but allocates no memory, and it simply aliases the shared buffer.
Heterogeneous Fast Path then runs that subgraph outside MLX's generic scheduler, without cross-device fences.

Metal's memory model is weak, so the authors don't depend on fine-grained coherence. No address is ever written by both devices, so there's no data race. The barrier supplies the one happens-before ordering the design needs.
On the CPU side, the SME kernels use 16×64 micro-tiles, multi-vector loads (svld1_x4), an active set kept within the 8 MB effective L2, and lock-free work-stealing. The work-stealing keeps slower E-cores from holding back the P-cores.
Profiling shows SME matmul performance rising in steps at multiples of 16 and dropping off after 16. The global draft budget is therefore fixed at M = 16 slots, shared across all agents.
Each agent gets a Historical Accepted Length (HAL): an EMA of the absolute number of tokens accepted per verification step. The authors chose absolute counts over acceptance rate because a shallow draft can have a perfect acceptance rate while producing very few tokens.
A_i(t) = γ·k_i(t) + (1−γ)·A_i(t−1), with γ = 0.3 and A_i(0) = 4.0 chosen by grid searchL_min slots, and the remaining slots are split by a temperature-scaled softmax over HAL, rounded down
Here it is as runnable Python. The paper doesn't give values for L_MIN or TAU, so I've flagged both as assumptions:
# hal.py: reimplementation of EdgeAgent Sec. 3.2.1 equations (not official code)
import math
from dataclasses import dataclass
L_TOTAL = 16 # paper: hardware-aligned global draft budget
GAMMA = 0.3 # paper: EMA decay factor
A0 = 4.0 # paper: initial HAL prior
L_MIN = 1 # ASSUMPTION: paper does not state its value
TAU = 1.0 # ASSUMPTION: paper does not state its value
@dataclass
class Agent:
name: str
hal: float = A0
suspended: bool = False
def update_hal(agent: Agent, accepted: int) -> None:
# A_i(t) = gamma * k_i(t) + (1 - gamma) * A_i(t-1)
agent.hal = GAMMA * accepted + (1 - GAMMA) * agent.hal
def allocate(agents: list[Agent]) -> dict[str, int]:
active = [a for a in agents if not a.suspended]
n = len(active)
if n == 0:
return {}
residual = L_TOTAL - n * L_MIN
assert residual >= 0, "more active agents than L_TOTAL / L_MIN"
z = sum(math.exp(a.hal / TAU) for a in active)
return {
a.name: L_MIN + math.floor(residual * math.exp(a.hal / TAU) / z)
for a in active
}At the end of each verification step, the scheduler checks every agent's state. An agent waiting on a tool moves to a suspended queue, and its slots go back to the pool for the allocator to hand out again. Its KV cache is frozen in place in shared memory. When the tool returns, the agent resumes without re-running prefill.

This adds tool stalls to the allocator above. The acceptance trace is made up for illustration:
# sim.py
from hal import Agent, allocate, update_hal
agents = [Agent("orchestrator"), Agent("coder"), Agent("tool_worker")]
# Toy acceptance counts per verification step (NOT from the paper)
trace = {
"orchestrator": [1, 2, 1, 1, 2, 1],
"coder": [6, 7, 8, 7, 8, 8],
"tool_worker": [5, 6, 6, 0, 0, 6],
}
tool_stall_steps = {3, 4} # tool_worker waits on an external tool
for t in range(6):
agents[2].suspended = t in tool_stall_steps # suspend-and-yield
budget = allocate(agents)
print(f"step {t}: {budget} used={sum(budget.values())}/16")
for a in agents:
if not a.suspended:
update_hal(a, min(trace[a.name][t], budget[a.name]))Output:
step 0: {'orchestrator': 5, 'coder': 5, 'tool_worker': 5} used=15/16
step 1: {'orchestrator': 2, 'coder': 6, 'tool_worker': 6} used=14/16
step 2: {'orchestrator': 1, 'coder': 7, 'tool_worker': 7} used=15/16
step 3: {'orchestrator': 1, 'coder': 14} used=15/16
step 4: {'orchestrator': 1, 'coder': 14} used=15/16
step 5: {'orchestrator': 1, 'coder': 11, 'tool_worker': 3} used=15/16The simulation shows three things the paper doesn't discuss:
tool_worker comes back, it competes using its pre-stall HAL. That works here, but if a tool result changes what the agent generates next (for example, from JSON to free text), its HAL will be out of date for a few steps.

You can't pip install EdgeAgent, but its conclusions apply now:
Set draft length per agent, not per process. Your orchestrator and your JSON-producing workers need different speculation settings. If your runtime has only one global setting, run them as separate model instances.
Make tool calls give up their compute. If your orchestration layer holds an inference slot open while waiting on a tool, you're paying for idle time. Questions about decomposition, like whether you need multiple agents at all, come first. When multi-agent setups help and when they hurt covers that tradeoff.
Don't assume CPU+GPU helps during decode. On UMA chips, adding a second processor to bandwidth-bound decode can make it slower (0.99x in the table above). Co-execution pays off in prefill. In decode, it only works when combined with speculation that raises arithmetic intensity.
Track absolute accepted tokens, not acceptance rate. Even without EdgeAgent, logging HAL per agent tells you which roles are worth speculating on.
L_min ≥ 1, more than 16 concurrent agents can't fit in the budget. The excerpt doesn't describe admission control.