toktape mascot: a mint-haired chibi in headphonestoktapethe record, not a recording of it
The toktape card for this run: 28.6 tok/s decode · 1 stream · Q3_K_M

0:00 / 0:00
Download the record

DeepSeek-V4.1-Flash Q3_K_M · 1 stream · 347 GB model on 72 GB of VRAM

DeepSeek-V4.1-Flash-Q3_K_M-00001-of-00009.gguf · 28.6 tok/s decode · 1 stream · Q3_K_M

midagedev

DeepSeek-V4.1-Flash Q3_K_M (engram Q8, token embeddings BF16), 9 shards, ~347 GB — far more than the 72 GB of VRAM on the RTX A6000 + RTX 3090, so most of the experts stream from 252 GB of DDR4. llama-server b96.

One stream, warm cache. The card flags the short prompt (35 tokens), so read the decode figure, not prefill.

decode
28.6tok/s
prefill
52.4tok/s
ttft p50
682ms
modelDeepSeek-V4.1-Flash-Q3_K_M-00001-of-00009.gguf
quantisationQ3_K_M
enginellama-server b96
hostlinux
gpusNVIDIA RTX A6000 / NVIDIA GeForce RTX 3090
streams1
prompt setnot in a comparison set
recorded2026-09-13T20:29:00+09:00
toktapedev

Details

The transcript, the card as text, and why each caveat fired — read out of the record in your browser, the same way Replay draws it.

! 1 caveatshort_prompt_for_prefill. The card that comes with this record spells each one out.