ad0bfb8146db14e15e0f57d5c47e9679972b1198
Observed KV cache size for qwen2.5:7b at fixed 16384 context varying 224MB-896MB between model loads with flash attention off. The larger figure pushes total memory needs just past the GPU's free VRAM, so some loads only offload 25/29 layers instead of 29/29 -- causing DJ-agent pick latency to jump from ~1s to multiple minutes. This GPU (6GB) has very little slack for this model/context combination even after freeing obico and stable-diffusion's VRAM reservations. Claude-Session: https://claude.ai/code/session_01L7Rwa6guD5wK8F8tWQwcJX
docker-infrastructure
Languages
Shell
58.2%
Python
28%
JavaScript
10.5%
HTML
2.3%
DIGITAL Command Language
0.6%
Other
0.4%