0:00 / 0:00
Qwen3.6-35B-A3B UD-Q4_K_XL · 1 stream
Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf · 140 tok/s decode · 1 stream · UD-Q4_K_XL
Qwen3.6-35B-A3B in UD-Q4_K_XL instead of Q6_K, one stream, ik_llama.cpp. The smaller quant is about 10% faster than the Q6_K single-stream rung on the same card.
decode
140tok/s
prefill
198tok/s
ttft p50
180ms
| model | Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf |
| quantisation | UD-Q4_K_XL |
| engine | ik_llama.cpp c10fbbcc |
| host | linux |
| gpus | NVIDIA RTX A6000 / NVIDIA GeForce RTX 3090 |
| streams | 1 |
| prompt set | not in a comparison set |
| recorded | 2026-09-15T22:58:28+09:00 |
| toktape | 0.2.3 |
Details
The transcript, the card as text, and why each caveat fired — read out of the record in your browser, the same way Replay draws it.
! 3 caveats —
placement_contradicted, short_prompt_for_prefill, recorded. The card that comes with this record spells each one out.