0:00 / 0:00
DeepSeek-V4.1-Flash Q3_K_M · Korean prose, no thinking
DeepSeek-V4.1-Flash-Q3_K_M-00001-of-00009.gguf · 20.3 tok/s decode · 1 stream · Q3_K_M
DeepSeek-V4.1-Flash Q3_K_M (engram Q8, token embeddings BF16), 9 shards, ~347 GB — far more than the 72 GB of VRAM on the RTX A6000 + RTX 3090, so most of the experts stream from 252 GB of DDR4. llama-server b96.
A Korean prose prompt with thinking off, one stream, cold cache. Recorded to see whether Hangul output changes the token rate; it does not, noticeably.
decode
20.3tok/s
prefill
30.4tok/s
ttft p50
4817ms
| model | DeepSeek-V4.1-Flash-Q3_K_M-00001-of-00009.gguf |
| quantisation | Q3_K_M |
| engine | llama-server b96 |
| host | linux |
| gpus | NVIDIA RTX A6000 / NVIDIA GeForce RTX 3090 |
| streams | 1 |
| prompt set | not in a comparison set |
| recorded | 2026-09-13T20:53:10+09:00 |
| toktape | dev |
Details
The transcript, the card as text, and why each caveat fired — read out of the record in your browser, the same way Replay draws it.
! 2 caveats —
cold_cache, client_disagrees_with_server. The card that comes with this record spells each one out.