0:00 / 0:00
llama.cpp · Qwen3.8 Flash Next UD-Q4_K_XL · RTX 3090 alone · coding review 538 in / 1500 out · -ncmoe 43, no draft
Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf · 35.5 tok/s decode · 1 stream · UD-Q4_K_XL
Recorded in one sitting beside bloomery on the same file, same prompt, greedy, thinking off. llama.cpp mainline 53ed051ce on the RTX 3090 alone (250 W cap): -ngl 99 -fa on -ncmoe 43 -t 32, no draft; the GPU list names both cards because the box has both. A clip tape, not a timed table.
decode
35.5tok/s
prefill
73.1tok/s
ttft p50
7439ms
| model | Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf |
| quantisation | UD-Q4_K_XL |
| size | 177B params · 10B active · 111.3 GB · MoE 512/10 |
| engine | llama-server b11157 |
| host | linux |
| gpus | NVIDIA GeForce RTX 3090 / NVIDIA RTX A6000 |
| streams | 1 |
| workload | 538 in / 1500 out · cache 0% |
| context | 8192 window · 1 slot |
| config | fa on · ngl 99 · offload partial |
| machine | 268.3 of 550 W · cold cache |
| prompt set | not in a comparison set |
| recorded | 2026-10-01T14:01:18Z |
| toktape | v0.6.0 |
Details
The transcript, the card as text, and why each caveat fired — read out of the record in your browser, the same way Replay draws it.
! 4 caveats —
cold_cache, recorded, machine_contended, conditions_changed. The card that comes with this record spells each one out.Page views are counted with Cloudflare Web Analytics: no cookies, no fingerprinting. The toktape binary itself sends nothing anywhere.
