Skip to content

[bug]: 6.14.1 regression with pytorch_cuda_alloc_conf: backend:cudaMallocAsync — Z-Image bf16 + LoRA takes ~30 min whenever the transformer is (re)loaded (VRAM overflows into Windows shared memory) #9597

Description

@cochbild

Is there an existing issue for this problem?

  • I have searched the existing issues

Related but not the same: #9147 / PR #9572 (Qwen-Image encoder VRAM accounting), #9257 (partial loading budget).

Operating system

Windows

GPU vendor

Nvidia (CUDA)

GPU model

RTX 5080

GPU VRAM

16GB

Version number

6.14.1

Browser

Chrome

Python dependencies

  • invokeai 6.14.1 (Launcher install), Python 3.12.9
  • torch 2.11.0+cu128, torchvision 0.26.0+cu128 (as installed by the launcher; note the cuda extra pins 2.7.1+cu128)
  • diffusers 0.40.0, transformers 5.5.4, accelerate 1.15.0
  • NVIDIA driver 610.74, Windows 11 Pro build 26200, 64 GB RAM

Background GPU usage is minimal: ~1.8–2.0 GB of VRAM in use by the desktop and other apps before each test, with nothing else using CUDA.

invokeai.yaml (memory settings unchanged since before the upgrade — the same file was in use for the ~32 s/image runs on 6.13.6 and 6.13.7):

pytorch_cuda_alloc_conf: backend:cudaMallocAsync
enable_partial_loading: true
max_cache_ram_gb: 32
device_working_mem_gb: 4

What happened

After upgrading from 6.13.7 to 6.14.1, Z-Image Turbo bf16 + a LoRA becomes extremely slow (~30 minutes for a 10-step 1024x1024 image) whenever the transformer has to be loaded into VRAM. That happens on the first run after startup, and after any prompt change, because running the text encoder evicts the transformer. Repeating the same prompt while the transformer is still resident takes ~45 s. Without the LoRA, or with the Q8 GGUF transformer, there's no slowdown.

During the slow run, dedicated VRAM is full (~15.6 / 16.3 GB) and Windows reports ~4.5 GB of shared GPU memory in use, i.e. the NVIDIA driver's sysmem fallback is paging GPU allocations to system RAM. The GPU shows 99% utilization but only ~75 W power draw (normal load on this card is well above that), so it is stalled waiting on PCIe transfers.

Same model, LoRA, settings and config on 6.13.x, from my gallery's metadata timestamps (time between consecutive images):

Version Image after a prompt change (median) Same prompt repeated (median)
6.13.0 40 s (n=17) 32 s (n=37)
6.13.6 34 s (n=46) 31 s (n=135)
6.13.7 32 s (n=43) 33 s (n=137)
6.14.1 ~32 min (reproduced) 43–50 s

Settings: z_image_turbo_bf16 single-file checkpoint (11.7 GB), Qwen3 encoder + VAE taken from the diffusers "Z-Image Turbo" main model (encoder 7.7 GB bf16), one LoRA at 0.4, 10 steps, CFG 1.5, euler, 1024x1024, seed variance enabled.

Debug log (run A, fresh server start, first generation) — the cache itself behaves as designed:

Loaded model '…:text_encoder' (Qwen3Model) onto cuda device #0 in 11.93s. Total model size: 7672.25MB, VRAM: 7672.25MB (100.0%)
... (z_image_denoise starts)
Before unloading: model_total=11740 MB, model_vram=0 MB (0.0% %), vram_available=3097 MB
Offloading unlocked models with goal of making room for 11739.56MB of VRAM.
Unloaded …:text_encoder from VRAM to free 7672 MB.
After unloading: model_total=11740 MB, model_vram=0 MB (0.0% %), vram_available=10745 MB
Loaded model '…:transformer' (ZImageTransformer2DModel) onto cuda device #0 in 7.63s. Total model size: 11739.56MB, VRAM: 10740.18MB (91.5%)
After loading: model_total=11740 MB, model_vram=10740 MB (91.5% %), vram_available=-1361 MB
  Compute Device (cuda)          Limit:  9392.5 MB, Used: 10753.7 MB (114.5%), Available: -1361.2 MB (-14.5%)
  CUDA Memory Allocated:         10753.7 MB

So torch reports ~10.75 GB allocated, but Windows shows the process using roughly 13.6 GB dedicated (excluding ~2 GB used by the desktop at idle) plus ~4.5 GB shared — about 7 GB more than torch accounts for, which matches the size of the text encoder that was just offloaded. Note also that loading 10.7 GB of transformer reduced driver-free memory by ~12.1 GB (vram_available went 10745 → -1361), and the compute-device "Limit" shrank from 10777 MB to 9392 MB during the load.

Run B (same prompt, immediately after cancelling run A) — driver reports zero free VRAM even after the cache frees memory:

Before unloading: model_total=11740 MB, model_vram=10740 MB (91.5% %), vram_available=-4096 MB
Unloaded models (if necessary): vram_bytes_freed=0.00MB
After unloading: ... vram_available=-4096 MB
Unloaded 4140.68MB from the model being locked (…:transformer).
Loaded model '…:transformer' ... VRAM: 6599.50MB (56.2%)
After loading: ... vram_available=-4096 MB
  CUDA Memory Allocated:         6613.4 MB

vram_available=-4096 with device_working_mem_gb: 4 means torch.cuda.mem_get_info() free == 0, and unloading 4.1 GB did not raise it. At idle afterwards, Windows still shows ~15.6 GB dedicated + ~3 GB shared for the process, with torch reporting ~6.6 GB allocated. Once in this state every generation is degraded until the server is restarted.

Controlled retest (fresh server start; "Clear Model Cache" before every run so the transformer always loads from scratch; runs that spilled were cancelled from the UI after ~30–60 s; 1024x1024, 10 steps):

Test Transformer LoRA (r128 @ 0.4) Text encoder Result Shared-memory spill GPU power
T1 bf16 on runs (new prompt) stalled 4.3 GB ~75 W
T2 bf16 on skipped (cached conditioning from T1) stalled 4.8 GB ~83 W
T3 bf16 off runs 39.2 s graph time none normal
T4 Q8 GGUF on skipped (cached from T1) 41.9 s graph time none normal

T2 shows the text encoder is not the trigger. The trigger is loading the bf16 transformer and then applying the LoRA. A prompt change only matters because loading the encoder evicts the transformer, which then has to be reloaded and re-patched. That's why same-prompt repeats with the transformer still resident are fast. In T1, T2 and T3 the cache ends up in the same state after loading the transformer (~91% in VRAM, vram_available ≈ -1.35 GB). Only the LoRA runs spill. Caveat on T4: the Q8 transformer fits fully in VRAM with ~3 GB of headroom, so T4 alone doesn't separate "sidecar patching" from "smaller model".

Earlier runs (same session, one changed word each run so the encoder must run):

Run Transformer LoRA (r128 @ 0.4) Result Shared-memory spill GPU power
C Z-Image Turbo Q8 GGUF on 29.5 s graph time none ~220–230 W
D z_image_turbo bf16 off 27.6 s graph time none ~210–220 W
E z_image_turbo bf16 on stalled (cancelled) 1.8 GB within seconds of denoise start ~80 W
A z_image_turbo bf16 on stalled (~30 min when allowed to finish) ~4.5 GB ~75 W

In D and E the cache does the same thing (encoder fully offloaded, transformer partially loaded to ~80–86%), so the extra VRAM appears only when the LoRA is applied to the bf16 transformer. With the GGUF transformer the LoRA goes through the sidecar path and there's no problem.

Allocator A/B — the regression disappears with PyTorch's default allocator. The only change was commenting out pytorch_cuda_alloc_conf: backend:cudaMallocAsync and restarting. Same model, LoRA, settings and session procedure:

Test Scenario Result Shared-memory spill
A1 fresh start, bf16 + LoRA, new prompt (= T1) 43.4 s (denoise 30.3 s, ~230 W) none
A2 prompt change with transformer resident 27.5 s none
A3 Clear Model Cache, same prompt (= T2) 34.3 s none
Series 3 prompts × 3 iterations, back to back first of each new prompt 35.6 s* / 27.0 s / 29.0 s; repeats 25–27 s none (peak 404 MB = desktop baseline)

* The first image in the series ran right after Clear Model Cache, so the encoder loaded from disk (10 s).

The cache's own numbers are the same as in the failing runs (transformer ~92% in VRAM, vram_available ≈ -1.8 GB after load), so the budgeting didn't change; only the allocator did. cudaMallocAsync worked fine with the same config on 6.13.0–6.13.7 (hundreds of images at ~32 s), so something in 6.14 interacts badly with it. The best candidate is the LoRA patching change below: allocating a fresh tensor per patched layer and freeing the old one leaves memory in the async allocator's pool, which is stream-ordered and may not be reused or returned before the next allocations.

Suspected cause: direct LoRA patching changed from in-place to out-of-place between 6.13.7 and 6.14.1 (invokeai/backend/patches/layer_patcher.py, _apply_model_layer_patch):

# 6.13.7
module_param.data.copy_(module_param.data + param_weight_converted)
# 6.14.1
module_param.data = module_param.data + param_weight_converted

The 6.14.1 comment says this is memory-equivalent, which holds only if the previous CUDA storage is freed as soon as .data is reassigned. If anything else still references the original GPU tensors (the partial-load bookkeeping, for example), each patched layer's weights are resident twice. That would explain several GB of VRAM the cache doesn't budget for, and the driver then satisfies it from system memory on Windows. (Inference — I haven't instrumented this.)

Other observations:

  • A standalone bf16 Qwen3 encoder (not from the diffusers folder) behaves identically (~4.4 GB spill).
  • The quantized GGUF Qwen3 encoder (3.3 GB) reduces the spill to ~0.6 GB and the image to ~1m54s, i.e. the spill scales with encoder size.
  • With the cached prompt after a completed slow run, the same-prompt image was ~50 s with no spill.

What you expected to happen

Same behaviour as 6.13.7: the text encoder is offloaded and its VRAM actually returned before the transformer is loaded, with no sysmem fallback, ~32 s per image regardless of prompt changes.

How to reproduce the problem

  1. Windows, single 16 GB NVIDIA GPU, config above (partial loading on, pytorch_cuda_alloc_conf: backend:cudaMallocAsync — required to reproduce).
  2. Start the server; generate with Z-Image Turbo bf16 + Qwen3 bf16 encoder + any Z-Image LoRA (mine is r128 at 0.4) at 1024x1024, 10 steps.
  3. Watch Task Manager → GPU → "Shared GPU memory" during denoise: several GB appear and denoise slows by ~40x.
  4. Or: Clear Model Cache, then generate again with the same prompt: same result, even though the encoder doesn't run.
  5. The same generation with the LoRA disabled completes in ~30–40 s with no shared-memory use.
  6. Remove the pytorch_cuda_alloc_conf line and restart: everything runs at ~25–35 s.

Additional context

  • Diffing 6.13.7 vs 6.14.1, the offload/VRAM-accounting paths (_load_locked_model, _offload_unlocked_models, _get_vram_available) look unchanged for single-GPU; the bigger changes on this path are from the multi-GPU work (feat: multi-GPU parallel session execution #9263: shared CPU weights store, global RAM budget, first-use holds/finalizers, deferred gc/empty_cache worker, LoRA patching by tensor replacement). I haven't identified which one is responsible.
  • Full debug logs for runs A–E and the T1–T4 retest, and nvidia-smi / Windows GPU memory counters sampled every 3–5 s, will be attached in a follow-up comment.
  • The Low-VRAM guide (https://invoke.ai/configuration/low-vram-mode/) currently recommends pytorch_cuda_alloc_conf: "backend:cudaMallocAsync", which is how this setting ended up in my config. Users who followed that advice will hit this regression on 6.14.x. Until it's fixed, the guide could use a caveat.
  • "Clear Model Cache" releases all of the spilled memory (VRAM returns to the ~1.8 GB desktop baseline), so the extra memory is held by live references rather than lost.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions