High-performance LoRA training on RTX 4090
Proved impossible: FLUX.2 Dev (~8B) Rank 1280 @ ~39s/it on a single RTX 4090 (24GB VRAM). Fully resolved PCIe bandwidth bottlenecks and VRAM overflows during extreme-rank LoRA training.
| Parameter | Standard / Default | Orakul Optimized |
|---|---|---|
| LoRA Rank | 128 / 256 | 1280 (Extreme Density) |
| Iteration Speed | ~65–70s / step | ~39–40s / step |
| Transformer Offload | 0.75 – 0.85 | 0.65 (PCIe Bottleneck Bypass) |
| Power Consumption | ~280–350W | ~136W (FP8 E5M2 Efficiency) |
| VRAM Stability | High Risk / OOM on Checkpoints | Zero OOM (Manual Latent Cache Eviction) |
Important
EN ⚡ CRITICAL PERFORMANCE NOTE: TRANSFORMER OFFLOAD RATIO
- Ranks from 32 to 1024 (Standard & High-Rank): Always set
offload = 0.75(Golden Ratio). This provides the ultimate balance between VRAM utilization and PCIe bus throughput (~6.5s/it). - Rank 1280 (Extreme 8B Parameter LoRA): Switch
offload = 0.65. Keeping an extra 10% of transformer blocks directly in VRAM circumvents the severe PCIe bus bottleneck for massive adapter matrices, unlocking peak training speed (~39s/it).
Do not use 0.65 for low/medium ranks, as it causes VRAM fragmentation and PCIe queueing.
Important
RU ⚡ КРИТИЧЕСКИ ВАЖНО: НАСТРОЙКА TRANSFORMER OFFLOAD RATIO
- Ранги от 32 до 1024 (Стандарт): Строго
0.75(Golden Ratio). Обеспечивает идеальный баланс VRAM и шины PCIe, отдавая стабильные ~6.5s на шаг. - Ранг 1280 (Экстремальный 8B LoRA): Переключайте на
0.65. Удержание ключевых слоёв базовой модели в VRAM полностью снимает пробки на шине PCIe при вычислении гигантских матриц адаптера, выбивая скорость ~39s на шаг.
Внимание: Не используйте 0.65 для стандартных рангов (128–512) — это вызовет фрагментацию VRAM и лишние микропростои шины.
- PCIe Bottleneck Bypass: Lowering
layer_offloading_transformer_percentto0.65kept critical matrices inside GDDR6X, removing GPU stall states. - Zero-OOM Latent Cache Clearing: Async RAM/VRAM cache clearing prevents memory leak spikes during step checkpoint saves.
Logs: HYPER-SPEED 17.09.2026
Configuration file: yaml
AI-Toolkit (Windows 11) is a user-friendly, high-performance toolkit designed for training diffusion models on consumer-grade hardware.
Built upon the open-source AI-Toolkit framework, this version has been extensively re-engineered to deliver server-grade training speeds and exceptional VRAM efficiency within the Windows 11 environment. Server-class speed on consumer hardware
Based on ostris/ai-toolkit - the original author's work is the foundation of everything here. All original commits preserved.
Why Orakul Studio is Windows Native Many people are used to thinking that "serious" development and AI belong on Linux. But if you look under the hood of any popular operating system, you'll see an endless series of workarounds, emulation attempts, and compromises.
We don't compromise.
Hardware Performance: The Viking Engine is optimized for Windows not out of laziness, but because it provides direct, low level access to the RTX 4090's resources without any slowdowns.
Death to workarounds: Linux is beautiful when you're browsing the web. But when it comes to working with heavyweights, asynchronous memory streaming, and real computing resources, it turns into a patchwork of patches. We value our time and studio resources more than "pretty" fonts in the terminal.
Unique Architecture: Our optimizations are the result of extensive work with the Ada Lovelace architecture. It works where it's supposed to fast, predictable, and without surprises like broken drivers or incompatible libraries.
Orakul Studio is designed for those who want to get things done, not just "administer" systems. Our base is Windows. Our goal is results. Anything else is just a waste of time. Linux is great for servers and clean code. But when you have two hours of daylight and need to run 200 steps to rank 1280, you don't care about a "proper kernel contract." You just want it to work.
So yes, my code is written for Windows. And it runs. If anyone wants to port the Viking Engine to Linux, they're welcome, I'm not opposed. But for now, I'll stick with where it's less of a hassle and more rewarding.
Contact and Support:
For questions regarding the proprietary memory management layer, integration nuances, or non-standard configurations, please email me at orakulstorm@gmail.com
Orakul Studio - Chernihiv, Ukraine 🇺🇦
If you're one of the 400+ people who cloned this repository, forget about any web UIs for this pipeline.
This code was designed, rewritten, and optimized exclusively for directly running configuration files (.yaml) via the console.
- Dynamic Alpha DOESN'T WORK AT ALL**
- This repository implements dynamic Alpha recalculation logic for correct weight scaling (Scale = Alpha / Rank). For example, when working with high ranks (Rank 128, Rank 512, Rank 1024), the system automatically calculates a fair scale (down to Scale = 0.5000), allowing the model to deeply learn the structure and physics of the material.
- The web UI completely ignores this logic. Almost all web wrappers under the hood forcibly overwrite this parameter and force a fixed Alpha = 16. At high ranks, this turns training into a dud: weight changes are suppressed, gradients tend to zero, the model visually "learns" without errors, but produces default output.
- Asynchronous Memory Manager (Async CUDA Memory Manager) CRASHED
- The logic for memory retention and low-level logging is optimized for the terminal's stdout.
- Web interfaces attempt to intercept and parse the string stream for their browser consoles. At best, this leads to a crash of the backend interface due to custom security prints; at worst, to a hidden downcast of tensor precision and gradient castration, so that a casual user doesn't simply "get a memory error."
- Run strictly through the console, directly from your virtual environment.
- If your process crashes while working with high ranks and dynamic alpha, don't look for compromises in the code; instead, increase the system swap/pagefile. The terminal works with your hardware without censorship or hidden precision reductions.
- Here is the configuration file for running the training Test1280.yaml
The original ai-toolkit is an excellent, flexible framework. This fork takes it in one specific direction: maximum performance on RTX 4090 (Ada Lovelace, sm_89).
Not about making weak hardware work. About making strong hardware fly.
| LoRA Configuration | Speed (s/it) | VRAM Memory Status | |
|---|---|---|---|
| Rank 128 (Optimized) | 6.70s / 6.50s | 24 GB (Zero OOM / Stable) | |
| Rank 512 (Deep Gesture) | 8.97s | 24 GB (Double Buffered) | |
| Rank 1024 (Extreme) | 22.45s | 24 GB (Full 8-bit Stack Forced) | |
| Rank 1280 (Extreme) | 65.80s | 24 GB (Full 8-bit Stack Forced) |
| Version | Speed | What Changed |
|---|---|---|
| Baseline (original) | 179 s/it | — |
| Viking v1 | 37 s/it | Double-buffer async CUDA |
| Viking v2 | 14 s/it | + bf16 weight forcing |
| Oracle-60 | 8.7 s/it | + Hardware FP8 + CPU prequant |
| LEGEND | 7.3 s/it | + Full 8-bit stack (AdamW 8-bit) |
24.5× faster than baseline. Same hardware. Zero OOM at rank 1280 (7.8B trainable params).
toolkit/manager_modules.pyd — Viking Engine
Double-buffered async weight streaming.
While GPU computes layer N, weights for layer N+1 transfer in a parallel CUDA stream. Transfer disappears from the profiler entirely.
> 🔒 **Orakul Studio Proprietary Tech**
> Core architecture and high-performance memory optimization layers are closed-source. Distributed exclusively via compiled binary module. The repository is open, and the pipeline is fully functional and stable..Also: CPU pinned memory for direct DMA from DRAM without CPU cache copy.
toolkit/quantize.py — Protocol Oracle-60
Native FP8 (E5M2) for Ada Lovelace + CPU pre-quantization.
RTX 4090 has native FP8 Tensor Cores. This activates them.
CPU pre-quantizes transformer blocks before GPU load PCIe bus freed.
"bf8": Float8WeightOnlyConfig(weight_dtype=torch.float8_e5m2)>>> [ORACLE-60] BF8 NATIVE (E5M2) DETECTED - CPU PRE-QUANT MODE
Why E5M2: preserves dynamic range like BF16. No "muddy faces" from aggressive quantization. Skin texture survives.
toolkit/lora_special.py Viking Override
LoRA matrices born in bfloat16, not converted later.
# === VIKING OVERRIDE: ЖЕСТКИЙ BFLOAT16 ДЛЯ ВЕСОВ ===
dtype = torch.bfloat16All Linear and Conv2d LoRA layers initialized directly in bf16. Half the memory at birth. PCIe transfer halved from step one.
Also fixed: Alpha/scale calculation for rank 1024.
Original code skipped .alpha keys during save, causing 64× signal drop at rank 1024 with alpha 64. Fixed and locked:
# === ORACLE STANDART PASS (FIXED) ===
alpha_val = alpha if alpha is not None and alpha != 0 else lora_dim
self.scale = alpha_val / self.lora_dimtoolkit/network_mixins.py — Oracle-60 Alpha Fix
The alpha skip bug that corrupted high-rank training.
Lines that skipped .alpha keys for non-LoKR networks are permanently commented with explanation. The lora_special.py autopilot handles scale correctly now.
# === [ORACLE-60] BLOCK: ЗАЩИТА ОТ МЫЛА ===
# These lines skipped Alpha for all types except LoKR.
# At Rank 1024, signal dropped 64x (Scale 16/1024 instead of 64/1024).
# KEEP DISABLED lora_special.py autopilot handles this correctly.toolkit/style.py — Perceptual Loss (Rewritten)
VGG19 perceptual loss without the VRAM leak.
Original had a critical bug: VGG19 was computing gradients for itself during perceptual loss calculation — wasting gigabytes of VRAM on a frozen reference network.
Fixed:
# Freeze VGG19 completely — it's a reference, not a trainee
for param in cnn.parameters():
param.requires_grad = FalseAlso: Gram matrix calculation replaced with hardware torch.bmm (Batch Matrix Multiply) — faster, cleaner, no nested function overhead.
Use perceptual loss for: Aivazovsky, watercolor, oil painting, charcoal, any artistic style where brushstroke texture matters more than pixel accuracy.
Use MSE for: portraits, photorealism, identity training.
toolkit/timer.py — Timer Silenced
CPU overhead from time polling eliminated.
Original timer called time.time() on every micro-step of the pipeline — predict_unet, backward, optimizer_step, etc. Parasitic CPU load, log spam, potential micro-freeze points on PCIe bus under heavy load.
def start(self, timer_name):
return # CPU freed
def stop(self, timer_name):
return # bus cleared
def print(self):
for hook in self._after_print_hooks:
hook({}) # engine hooks get empty dict, nothing breaks
returnResult: cleaner logs, lower CPU load, stable PCIe bus under sustained training.
jobs/process/BaseSDTrainProcess.py — bf16 Forcing
Network forced to bfloat16 before training starts.
# Viking method — before network.apply_to()
# todo switch everything to proper mixed precision like this ← ostris left this todo
self.network.force_to(self.device_torch, dtype=torch.bfloat16)The # todo comment was already in the original source. We read it and implemented it.
For portraits, faces, identity training. Pixel-accurate, fast.
content_or_style: balanced
loss_type: mseFor Aivazovsky, watercolor, oil, charcoal. Learns brushstroke, not pixels.
Uses ~1-2 GB more VRAM, ~20-30% slower, higher GPU voltage.
content_or_style: style
loss_type: perceptualSwitch between modes by commenting/uncommenting — no config rewrite needed.
network:
type: lora
linear: 1280 # rank 1280 = 7.8B trainable params
linear_alpha: 64
conv: 32
conv_alpha: 64
lokr_full_rank: true
lokr_factor: -1
model:
qtype: bf8 # Oracle-60 native FP8
quantize_te: true
qtype_te: bf8
layer_offloading: true
layer_offloading_transformer_percent: 0.91Web UI works but sampling is disabled for maximum speed.
For real performance — terminal only:
# Activate environment
source venv/Scripts/activate
# Run training with full log capture
python run.py viking_train/your_config.yaml 2>&1 | tee "log_training.txt"Configs go in viking_train/ folder.
- RTX 4090 (24 GB) — optimized for Ada Lovelace sm_89
- 64+ GB RAM recommended (128 GB for rank 1280)
- Python 3.10+ / Linux or Windows 11
git clone https://github.com/OrakulStudio/ai-toolkit.git
cd ai-toolkit
python -m venv venv
.\venv\Scripts\activate
pip install --no-cache-dir torch==2.9.1 torchvision==0.24.1 torchaudio==2.9.1 --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt
If you encounter numpy.dtype errors or CUDA/Triton warnings, use the following commands to ensure your environment is set up correctly for high-performance training:
Fix NumPy/SciPy version conflicts:
pip install "numpy<2.0.0" scipy --force-reinstall
---
Configs go in viking_train/ folder.
What Ostris Built
This fork exists because ostris built something worth building on.
ostris/ai-toolkit is the most widely used open-source LoRA training framework. Thousands of people use it daily. It's clean, flexible, actively maintained.
All original commits preserved. Author credited. The original README can be found in README_OSTRIS.md.
---
Image
- black-forest-labs/FLUX.1-dev (FLUX.1)
- black-forest-labs/FLUX.2-dev (FLUX.2)
- black-forest-labs/FLUX.2-klein-base-4B (FLUX.2-klein-base-4B)
- black-forest-labs/FLUX.2-klein-base-9B (FLUX.2-klein-base-9B)
- stabilityai/stable-diffusion-xl-base-1.0 (SDXL)
- HiDream-ai/HiDream-I1-Full (HiDream I1)
- Qwen/Qwen-Image (Qwen-Image)
- Tongyi-MAI/Z-Image (Z-Image)
- Tongyi-MAI/Z-Image-Turbo (Z-Image Turbo)
- lodestones/Chroma1-Base (Chroma)
- OmniGen2/OmniGen2 (OmniGen2)
Instruction / Edit
- black-forest-labs/FLUX.1-Kontext-dev (FLUX.1-Kontext-dev)
- Qwen/Qwen-Image-Edit (Qwen-Image-Edit)
- HiDream-ai/HiDream-E1-1 (HiDream E1)
Video
- Wan-AI/Wan2.2-T2V-A14B-Diffusers (Wan 2.2 14B)
- Wan-AI/Wan2.2-I2V-A14B-Diffusers (Wan 2.2 I2V 14B)
- Wan-AI/Wan2.2-TI2V-5B-Diffusers (Wan 2.2 TI2V 5B)
- Lightricks/LTX-2 (LTX-2)
- 🐙 Original: ostris/ai-toolkit
- 🐙 This fork: github.com/OrakulStudio
- 🤗 huggingface.co/OrakulStorm
- 🎨 civitai.com/user/orakul_studio
*The smell of the iron is stable. 🦊⚡*
*Chernihiv, Ukraine 🇺🇦 · Orakul Studio · 2026*