0:00 / 0:00
Qwen3.8-Flash-Next UD-Q4_K_XL · 1 stream · mlock
Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf · 51.0 tok/s decode · 1 stream · UD-Q4_K_XL
Qwen3.8-Flash-Next in UD-Q4_K_XL, one stream, llama-server, weights mlocked. Fourth take of a loading-mode comparison (mmap vs mlock vs none): mlock kept the run warm from the first token.
decode
51.0tok/s
prefill
322tok/s
ttft p50
1411ms
| model | Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf |
| quantisation | UD-Q4_K_XL |
| engine | llama-server b0 |
| host | linux |
| gpus | NVIDIA GeForce RTX 3090 / NVIDIA RTX A6000 |
| streams | 1 |
| prompt set | not in a comparison set |
| recorded | 2026-09-16T10:09:12+09:00 |
| toktape | 0.2.3-5-g9bf4e52 |
Details
The transcript, the card as text, and why each caveat fired — read out of the record in your browser, the same way Replay draws it.
