0:00 / 0:00
llama.cpp · Qwen3.8 Flash Next UD-Q4_K_XL · A6000 + 3090 · coding review 538 in / 1500 out · -ncmoe 21, no draft
Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf · 44.3 tok/s decode · 1 stream · UD-Q4_K_XL
Recorded in one sitting beside bloomery on the same file, same prompt, greedy, thinking off. llama.cpp mainline 53ed051ce on the A6000 + 3090: -ngl 99 -fa on -ncmoe 21 -ts 42.5/6.5 -t 32, no draft. A clip tape, not a timed table.
decode
44.3tok/s
prefill
217tok/s
ttft p50
2583ms
| model | Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf |
| quantisation | UD-Q4_K_XL |
| size | 177B params · 10B active · 111.3 GB · MoE 512/10 |
| engine | llama-server b11157 |
| host | linux |
| gpus | NVIDIA GeForce RTX 3090 / NVIDIA RTX A6000 |
| streams | 1 |
| workload | 538 in / 1500 out · cache 0% |
| context | 8192 window · 1 slot |
| config | fa on · ngl 99 · offload partial |
| machine | 373.9 of 550 W · cold cache |
| prompt set | not in a comparison set |
| recorded | 2026-10-01T13:46:26Z |
| toktape | v0.6.0 |
Details
The transcript, the card as text, and why each caveat fired — read out of the record in your browser, the same way Replay draws it.
! 3 caveats —
cold_cache, machine_contended, conditions_changed. The card that comes with this record spells each one out.Page views are counted with Cloudflare Web Analytics: no cookies, no fingerprinting. The toktape binary itself sends nothing anywhere.
