obico's own comment already notes its :cuda tag falls back to CPU inference on this driver/GPU combo, but it was still reserving 826MB of VRAM it never used productively. That squeeze was forcing ollama to offload only 20/29 model layers to GPU, pushing the rest onto CPU and causing severe latency (multi-minute LLM calls) that stalled SUB/WAVE's track picking. Claude-Session: https://claude.ai/code/session_01L7Rwa6guD5wK8F8tWQwcJX