DeepSeek-V4.1-Flash Q3_K_M · Korean prose, no thinking 20.3 tok/s DeepSeek-V4.1-Flash-Q3_K_M-00001-of-00009.gguf llama-serverb96Q3_K_Mrtx-a6000rtx-3090linux1 stream! 2 caveats midagedev 2026-09-18
DeepSeek-V4.1-Flash Q3_K_M · 4 streams 27.4 tok/s DeepSeek-V4.1-Flash-Q3_K_M-00001-of-00009.gguf llama-serverb96Q3_K_Mrtx-a6000rtx-3090linux4 streams! 1 caveat midagedev 2026-09-18
DeepSeek-V4.1-Flash Q3_K_M · 1 stream · 347 GB model on 72 GB of VRAM 28.6 tok/s DeepSeek-V4.1-Flash-Q3_K_M-00001-of-00009.gguf llama-serverb96Q3_K_Mrtx-a6000rtx-3090linux1 stream! 1 caveat midagedev 2026-09-18