0:00 / 0:00
GLM-5.3-Flash EXL3 4.05 bpw · 2 streams · exllamav3
GLM-5.3-Flash-exl3-4.05 · 19.4 tok/s decode · 2 streams · EXL3 4.05 bpw · head 6.0
Same server, two streams batched. Per-stream decode halves and the aggregate lands near the single-stream rate — batching on exllamav3 did not add throughput for this model on this pair of cards.
decode
19.4tok/s
prefill
78.3tok/s
ttft p50
2780ms
| model | GLM-5.3-Flash-exl3-4.05 |
| quantisation | EXL3 4.05 bpw · head 6.0 |
| engine | exllamav3 1.5.0 |
| host | linux |
| gpus | NVIDIA RTX A6000 / NVIDIA GeForce RTX 3090 |
| streams | 2 |
| prompt set | not in a comparison set |
| recorded | 2026-09-15T22:26:36+09:00 |
| toktape | 0.2.2 |
Details
The transcript, the card as text, and why each caveat fired — read out of the record in your browser, the same way Replay draws it.
! 2 caveats —
recorded, conditions_changed. The card that comes with this record spells each one out.