0:00 / 0:00
bloomery · GLM-5.3-Flash UD-Q4_K_XL · A6000 alone · Korean chat 42 in / 705 out · residency + NextN MTP draft
GLM-5.3-Flash-UD-Q4_K_XL-00001-of-00006.gguf · 30.8 tok/s decode · 1 stream · q4_K
GLM-5.3-Flash UD-Q4_K_XL (public file), bloomery 07ec8f04 bloomery-serve --model glm --place a --ctx 4096 on the RTX A6000 alone, the serving defaults: adaptive residency mid-p0-s1 + the NextN MTP draft, prompts in groups of two 512-position batches with the host union's row lanes; greedy, cache_prompt off, reasoning_effort low (GLM-5.3's template always thinks).
decode
30.8tok/s
prefill
75.5tok/s
ttft p50
619ms
| model | GLM-5.3-Flash-UD-Q4_K_XL-00001-of-00006.gguf |
| quantisation | q4_K · 4.5 bit |
| size | 199.7 GB · MoE 288/8 |
| engine | bloomery 0.1.0 (07ec8f04) |
| host | linux |
| gpus | NVIDIA GeForce RTX 3090 / NVIDIA RTX A6000 |
| streams | 1 |
| workload | 42 in / 705 out · cache 0% |
| context | 4096 window · 1 slot |
| config | offload partial |
| draft | GLM-5.3-Flash-UD-Q4_K_XL-00001-of-00006.gguf · 82% accepted |
| machine | 328.7 of 550 W |
| prompt set | not in a comparison set |
| recorded | 2026-10-02T23:14:39Z |
| toktape | v0.6.0 |
Details
The transcript, the card as text, and why each caveat fired — read out of the record in your browser, the same way Replay draws it.
! 3 caveats —
short_prompt_for_prefill, client_disagrees_with_server, conditions_changed. The card that comes with this record spells each one out.Page views are counted with Cloudflare Web Analytics: no cookies, no fingerprinting. The toktape binary itself sends nothing anywhere.
