Selection of the highest-performance setup. Correctness is an eliminatory gate; among the correct ones, warm total time, TTFT and sustained decode decide.
canonical = vendor sampling (decision metric); greedy = diagnostic temp=0, dimmed, does not count for verdict and inflates decode. E2E does not compare across classes (code short vs audit long). Arms with cache off (W/X) have cold TTFT by design.
■ canonical (temp=1, counts for verdict) · ■ greedy only (temp=0, diagnostic) · — not run. The tool loop column shows majority (passes/total). Not every arm runs all scenarios — depends on the gate.
T promoted (M4 production); U just control, no advantage >5%.
Best decode 32K; tool loop 2/3 with corrected prompt.
DFlash2 +14-24% decode, but no 10% of total time (cache-off).
L beats K on total time; MTP acceptance ~0.8; tool loop 2/3.
Cuts cold TTFT ~-55%, but zeroes the warm cache (hit 0%); warm TTFT 6-17x worse than L. None advances.
Blocked: 'Private ANE procedure-bank compiler is unavailable' in this omlx build. ANE kernels do not compile. External, not fixable in the campaign.
Speculation delivers: decode +90%/+154% and total time -57%/-70% vs Q; tool loop 2/3. But the peak (S 38 @32K) falls below MTPLX (44.5) and is 8-bit.
Speed ~23-31% faster in decode. Among Quality, Z(fp16) dominates Y(8-bit): same decode, tool loop 3/5 vs 1/5. Quality only for fidelity (not measured).
At the native maximum (262K) the MTPLX MTP collapses: decode V 6.1 / Y 8.7 tps. oMLX (L 15.4 / T 14.1) and dspark/DFlash2 (S 14.2) hold ~14-15 tps. Cause: the MTP verify re-reads the 262K KV per step; DFlash2 does not scale that way. Speed verdict at the ceiling: oMLX and dspark; MTPLX loses the advantage it had up to 128K.
All reuse the prefix at the ceiling: identical ~1.0, append 0.95-1.0, tool_turn 0.99-1.0; middle_mutation partial (0.38-0.49). CONFIRMED by variable isolation: the MTPLX tool_turn that zeroed at 128K/cap 48G reuses 0.99 at 128K/cap 100G (same worst-case order) -> the miss was session-bank under-provisioning, not order. Rule: size the per-session cap for the context KV + tool footprint (48G is too little already at 128K; 100G suffices at 262K).
canonical = temp=1 (counts for verdict); greedy = temp=0 (diagnostic). — in context = planned arm, not yet run. MTP: ✓ native, draft via draft model, auto vendor default.