0:00 / 0:00
DeepSeek-V2-Lite-Chat Q3_K_M · 4 streams
DeepSeek-V2-Lite-Chat.Q3_K_M.gguf · 211 tok/s decode · 4 streams · Q3_K_M · rtx-a6000
The same model at four streams. Per-stream rate falls to about a third; the aggregate is a little over 200 tok/s.
decode
211tok/s
prefill
2577tok/s
ttft p50
299ms
| model | DeepSeek-V2-Lite-Chat.Q3_K_M.gguf |
| quantisation | Q3_K_M |
| engine | ik_llama.cpp c10fbbcc |
| host | linux · 1x rtx-a6000 |
| gpus | NVIDIA RTX A6000 |
| streams | 4 |
| prompt set | not in a comparison set |
| recorded | 2026-09-17T18:32:07+09:00 |
| toktape | 0.2.2 |
Details
The transcript, the card as text, and why each caveat fired — read out of the record in your browser, the same way Replay draws it.
! 1 caveat —
recorded. The card that comes with this record spells each one out.