DeepSeek-V4.1-Flash Q3_K_M · Korean prose, no thinking 20.3 tok/s DeepSeek-V4.1-Flash-Q3_K_M-00001-of-00009.gguf llama-serverb96Q3_K_Mrtx-a6000rtx-3090linux1 stream! 2 caveats midagedev 2026-09-18
DeepSeek-V4.1-Flash Q3_K_M · 1 stream · 347 GB model on 72 GB of VRAM 28.6 tok/s DeepSeek-V4.1-Flash-Q3_K_M-00001-of-00009.gguf llama-serverb96Q3_K_Mrtx-a6000rtx-3090linux1 stream! 1 caveat midagedev 2026-09-18
DeepSeek-V2-Lite-Chat Q3_K_M · 1 stream 183 tok/s DeepSeek-V2-Lite-Chat.Q3_K_M.gguf ik_llama.cppc10fbbccQ3_K_Mrtx-a6000linux1 stream! 1 caveat midagedev 2026-09-18
Qwen2.5-7B-Instruct Q3_K_M · 1 stream · dense control 116 tok/s Qwen2.5-7B-Instruct-Q3_K_M.gguf ik_llama.cppc10fbbccQ3_K_Mrtx-a6000linux1 stream! 2 caveats midagedev 2026-09-18