Enable flash attention + quantized KV cache for ollama

Observed KV cache size for qwen2.5:7b at fixed 16384 context varying
224MB-896MB between model loads with flash attention off. The larger
figure pushes total memory needs just past the GPU's free VRAM, so
some loads only offload 25/29 layers instead of 29/29 -- causing
DJ-agent pick latency to jump from ~1s to multiple minutes. This GPU
(6GB) has very little slack for this model/context combination even
after freeing obico and stable-diffusion's VRAM reservations.

Claude-Session: https://claude.ai/code/session_01L7Rwa6guD5wK8F8tWQwcJX
This commit is contained in:
2026-09-13 15:30:36 +00:00
parent cd127a23e0
commit ad0bfb8146
+7
View File
@@ -10,6 +10,13 @@ services:
- NVIDIA_DRIVER_CAPABILITIES=compute,utility
- TZ=America/New_York
- OLLAMA_HOST=0.0.0.0
# This GPU (6GB) is right at the edge for qwen2.5:7b at 16384 context —
# KV cache size was observed varying 224MB-896MB between loads with flash
# attention off, and the larger figure tips offload to 25/29 layers
# instead of 29/29, causing multi-minute DJ-agent latency. Flash
# attention + an explicitly quantized KV cache keep it small and stable.
- OLLAMA_FLASH_ATTENTION=1
- OLLAMA_KV_CACHE_TYPE=q8_0
volumes:
- /srv/ollama:/root/.ollama
runtime: nvidia