cuda: bound the weight-cache arena to free VRAM by default - #995
Open
GHTKFI wants to merge 1 commit into
Open
Conversation
cuda_model_cache_limit_bytes() returned UINT64_MAX whenever DS4_CUDA_WEIGHT_CACHE_LIMIT_GB was unset, letting the weight-cache arena grow unbounded during prefill until a later cudaMalloc call fails mid-generation (issue antirez#170 shows the identical "arena alloc failed: out of memory" symptom on an RTX 4090, with the same unbounded default as the unidentified root cause). Query cudaMemGetInfo() and default the limit to free VRAM minus a 4 GiB safety margin instead, leaving room for the KV cache, prefill buffers, and the dynamic q8/fp16 cache. DS4_CUDA_WEIGHT_CACHE_LIMIT_GB still overrides this when set explicitly. Tested: reproduced the OOM crash on an RTX 3090 (24 GB) with a 22-token prompt + --ssd-streaming-cache-experts 14GB; with this fix, the same command completes generation with no arena alloc failures.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
cuda_model_cache_limit_bytes() returned UINT64_MAX whenever DS4_CUDA_WEIGHT_CACHE_LIMIT_GB was unset, letting the weight-cache arena grow unbounded during prefill until a later cudaMalloc call fails mid-generation (issue #170 shows the identical "arena alloc failed: out of memory" symptom on an RTX 4090, with the same unbounded default as the unidentified root cause).
Query cudaMemGetInfo() and default the limit to free VRAM minus a 4 GiB safety margin instead, leaving room for the KV cache, prefill buffers, and the dynamic q8/fp16 cache. DS4_CUDA_WEIGHT_CACHE_LIMIT_GB still overrides this when set explicitly.
Tested: reproduced the OOM crash on an RTX 3090 (24 GB) with a 22-token prompt + --ssd-streaming-cache-experts 14GB; with this fix, the same command completes generation with no arena alloc failures.