Qwen3.8-Flash-Next UD-Q4_K_XL · 1 stream · mlock 51.0 tok/s Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf llama-serverb0UD-Q4_K_XLrtx-3090rtx-a6000linux1 stream midagedev 2026-09-18
DeepSeek-V4.1-Flash Q3_K_M · Korean prose, no thinking 20.3 tok/s DeepSeek-V4.1-Flash-Q3_K_M-00001-of-00009.gguf llama-serverb96Q3_K_Mrtx-a6000rtx-3090linux1 stream! 2 caveats midagedev 2026-09-18
DeepSeek-V4.1-Flash Q3_K_M · 1 stream · 347 GB model on 72 GB of VRAM 28.6 tok/s DeepSeek-V4.1-Flash-Q3_K_M-00001-of-00009.gguf llama-serverb96Q3_K_Mrtx-a6000rtx-3090linux1 stream! 1 caveat midagedev 2026-09-18