0:00 / 0:00
DeepSeek-V4.1-Flash Q3_K_M · 1 stream · 347 GB model on 72 GB of VRAM
DeepSeek-V4.1-Flash-Q3_K_M-00001-of-00009.gguf · 23.8 tok/s decode · 1 stream · Q3_K_M
decode
23.8tok/s
prefill
52.3tok/s
ttft p50
124452ms
| model | DeepSeek-V4.1-Flash-Q3_K_M-00001-of-00009.gguf |
| quantisation | Q3_K_M |
| size | 748B params · 18B active · 477.0 GB · MoE 384/6 |
| engine | llama-server b102 |
| host | linux |
| gpus | NVIDIA GeForce RTX 3090 / NVIDIA RTX A6000 |
| streams | 1 |
| workload | 6509 in / 66 out · cache 0% |
| context | 8192 window · 4 slots |
| config | b 2048 · ub 512 · ngl 99 · offload partial |
| draft | DeepSeek-V4.1-Flash-Fp8-128x742M-MXFP4_MOE.tl37.gguf · 65% accepted |
| machine | 296.2 of 720 W · cold cache |
| prompt set | not in a comparison set |
| recorded | 2026-09-20T07:38:08Z |
| toktape | dev |
Details
The transcript, the card as text, and why each caveat fired — read out of the record in your browser, the same way Replay draws it.
! 4 caveats —
cold_cache, machine_contended, conditions_changed, run_cut_by_clock. The card that comes with this record spells each one out.