toktape mascot: a mint-haired chibi in headphonestoktapethe record, not a recording of it
The toktape card for this run: 21.5 tok/s decode · 1 stream · UD-Q4_K_XL

0:00 / 0:00
Download the record

llama.cpp · GLM-5.3-Flash UD-Q4_K_XL · A6000 alone · Korean chat 42 in / 674 out · PR #27754 MTP draft, -ncmoe 38

GLM-5.3-Flash-UD-Q4_K_XL-00001-of-00006.gguf · 21.5 tok/s decode · 1 stream · UD-Q4_K_XL

midagedev

llama.cpp PR #27754 llama-server (86ebfef2c) on the A6000 alone: -ngl 999 --n-cpu-moe 38 -fa off -t 32 -fit off --no-op-offload -np 1 -ctxcp 0 --cache-ram 0 --spec-type draft-mtp --spec-draft-n-max 2, NVIDIA_TF32_OVERRIDE=0, -c 4096. --n-cpu-moe 37 did not fit the MTP context's compute buffer at -c 4096. Shards preheated, two warm-ups discarded, greedy, cache_prompt off, reasoning_effort low through chat_template_kwargs.

decode
21.5tok/s
prefill
62.5tok/s
ttft p50
803ms
modelGLM-5.3-Flash-UD-Q4_K_XL-00001-of-00006.gguf
quantisationUD-Q4_K_XL
size321B params · 23B active · 199.7 GB · MoE 288/8
enginellama-server b11036
hostlinux
gpusNVIDIA GeForce RTX 3090 / NVIDIA RTX A6000
streams1
workload42 in / 674 out · cache 0%
context4096 window · 1 slot
configfa off · ngl 999 · offload partial
machine173.8 of 550 W
prompt setnot in a comparison set
recorded2026-10-01T18:55:37Z
toktapev0.6.0

Details

The transcript, the card as text, and why each caveat fired — read out of the record in your browser, the same way Replay draws it.

! 4 caveats — short_prompt_for_prefill, client_disagrees_with_server, recorded, conditions_changed. The card that comes with this record spells each one out.

Page views are counted with Cloudflare Web Analytics: no cookies, no fingerprinting. The toktape binary itself sends nothing anywhere.