Enable flash attention + quantized KV cache for ollama
Observed KV cache size for qwen2.5:7b at fixed 16384 context varying 224MB-896MB between model loads with flash attention off. The larger figure pushes total memory needs just past the GPU's free VRAM, so some loads only offload 25/29 layers instead of 29/29 -- causing DJ-agent pick latency to jump from ~1s to multiple minutes. This GPU (6GB) has very little slack for this model/context combination even after freeing obico and stable-diffusion's VRAM reservations. Claude-Session: https://claude.ai/code/session_01L7Rwa6guD5wK8F8tWQwcJX
This commit is contained in:
@@ -10,6 +10,13 @@ services:
|
|||||||
- NVIDIA_DRIVER_CAPABILITIES=compute,utility
|
- NVIDIA_DRIVER_CAPABILITIES=compute,utility
|
||||||
- TZ=America/New_York
|
- TZ=America/New_York
|
||||||
- OLLAMA_HOST=0.0.0.0
|
- OLLAMA_HOST=0.0.0.0
|
||||||
|
# This GPU (6GB) is right at the edge for qwen2.5:7b at 16384 context —
|
||||||
|
# KV cache size was observed varying 224MB-896MB between loads with flash
|
||||||
|
# attention off, and the larger figure tips offload to 25/29 layers
|
||||||
|
# instead of 29/29, causing multi-minute DJ-agent latency. Flash
|
||||||
|
# attention + an explicitly quantized KV cache keep it small and stable.
|
||||||
|
- OLLAMA_FLASH_ATTENTION=1
|
||||||
|
- OLLAMA_KV_CACHE_TYPE=q8_0
|
||||||
volumes:
|
volumes:
|
||||||
- /srv/ollama:/root/.ollama
|
- /srv/ollama:/root/.ollama
|
||||||
runtime: nvidia
|
runtime: nvidia
|
||||||
|
|||||||
Reference in New Issue
Block a user