0:00 / 0:00
GLM-5.3-Flash graftB Q6_K · MTP · ik_llama.cpp
GLM-5.3-Flash-graftB.gguf · 24.2 tok/s decode · 1 stream · Q6_K
A 184 GiB Q6_K GGUF of GLM-5.3-Flash (graftB) with MTP draft decoding on ik_llama.cpp 425a2c1d, across the A6000 + 3090 with the rest in RAM. One stream; the card flags the short prompt.
decode
24.2tok/s
prefill
49.8tok/s
ttft p50
768ms
| model | GLM-5.3-Flash-graftB.gguf |
| quantisation | Q6_K · 6.56 bit |
| engine | ik_llama.cpp 425a2c1d |
| host | linux |
| gpus | NVIDIA RTX A6000 / NVIDIA GeForce RTX 3090 |
| streams | 1 |
| prompt set | not in a comparison set |
| recorded | 2026-09-15T13:08:01+09:00 |
| toktape | dev |
Details
The transcript, the card as text, and why each caveat fired — read out of the record in your browser, the same way Replay draws it.
! 2 caveats —
short_prompt_for_prefill, recorded. The card that comes with this record spells each one out.