0:00 / 0:00
Qwen3.6-35B-A3B Q6_K · 8 streams on 4 slots · ik_llama.cpp
Qwen3.6-35B-A3B-UD-Q6_K.gguf · 152 tok/s decode · 8 streams · UD-Q6_K · rtx-a6000
Qwen3.6-35B-A3B UD-Q6_K (27.3 GiB) on one RTX A6000 48G, ik_llama.cpp c10fbbcc, -fa on.
Part of a 1/2/4/8-stream sweep on the same server and the same prompt set; the other rungs are published beside this one. Eight streams against four slots: half of them queue, which is what the 12 s TTFT is. The aggregate does not rise above the 4-stream rung — more concurrency than slots buys waiting, not throughput.
decode
152tok/s
prefill
146tok/s
ttft p50
12529ms
| model | Qwen3.6-35B-A3B-UD-Q6_K.gguf |
| quantisation | UD-Q6_K · 6.56 bit |
| engine | ik_llama.cpp c10fbbcc |
| host | linux · 1x rtx-a6000 |
| gpus | NVIDIA RTX A6000 |
| streams | 8 |
| prompt set | not in a comparison set |
| recorded | 2026-09-17T15:10:55+09:00 |
| toktape | 0.2.2 |
Details
The transcript, the card as text, and why each caveat fired — read out of the record in your browser, the same way Replay draws it.
! 2 caveats —
answer_cut, recorded. The card that comes with this record spells each one out.