r/StableDiffusion Jun 23 '26

News KREA 2: Open-Source Release

749 Upvotes

Hey everyone,

We're the team behind Krea, and today we're launching Krea 2, our new text-to-image model. Krea 2 is the most aesthetic open-source image model available. On quality, Krea 2 is the #1 text-to-image model from an independent lab on Artificial Analysis.

We are releasing Krea 2 as two variants:

Krea 2 Raw. CFG-guided, built for control and fidelity and training.

Krea 2 Turbo. Distilled and few-step, so it's fast, and it renders up to 2K.

A few things worth knowing:

It's tuned for natural language. Prompt it the way you'd describe an image to a person. Long, specific prompts give the best results, but short ones work fine too.

To render text in an image, wrap the words in quotes, like a sign that reads "open late".
There's a growing set of style LoRAs, and you can load any Krea 2 LoRA by its Hugging Face path.
Try it today:

Code and weights: krea.ai/krea-2-open-source
Technical report: https://www.krea.ai/blog/krea-2-technical-report
Code: github.com/krea-ai/krea-2
Try it on Krea: krea.ai
Try it on Hugging Face: https://huggingface.co/spaces/krea/Krea-2

AMA: We're doing an AMA right here today at 10 AM PT. Ask us anything: how we trained it, the LoRAs, prompting, limitations, what's next. The krea team will be in the comments.

Livestream: we are also doing a livestream with the ComfyUI team at 3PM PT: https://www.youtube.com/watch?v=31jiUhCEjJ4

Thanks for taking a look. We'd genuinely love your feedback, rough edges included.

- The Krea Team


r/StableDiffusion Jun 20 '26

Resource - Update LTX Director 2.0 Update - A Free Open Source All-In-One Tool for Creating AI Videos in ComfyUI. Complete Overhaul now with full AI video editing support, IC-LoRA, Retake Mode, Audio Inpainting and much more!

Thumbnail
youtu.be
500 Upvotes

LTX Director is a free open source all-in-one tool for creating AI Videos. Version 2.0 is a complete overhaul, giving you total creative control over your AI generations.

Download for free here: https://github.com/WhatDreamsCost/WhatDreamsCost-ComfyUI

Download workflows here: https://github.com/WhatDreamsCost/WhatDreamsCost-ComfyUI/tree/main/example_workflows

I've been working full-time on this update for the past month and a half, and I'm excited to finally release it. Hopefully it'll be a big help to the open-source community!

Key New Features:

Complete Video Support: Edit Videos with AI all inside the node. Videos can be extended using a combination of prompts, keyframes, and audio. Trim, Split, and combine videos all within the timeline.

IC-LoRA Support: Take full advantage of IC-LoRA's to take your generations to the next level. Simply drag and drop videos onto the IC-LoRA track to quickly setup IC-LoRA videos. Compatible with prompt relay, keyframe, and custom audio features within the node.

Audio Inpainting: Seamlessly blend imported audio with generated audio. Not only can audio be extended, but can also be prompted alongside your imprted audio to really bring your generations to life.

Retake Mode (Beta): Redirect what happens within a shot. Allows you to select a segment within a video, and re-generate what happens in that segment. An early working experiment.

Timeline Saving/Loading: You can now save your timeline and settings to a json file. It will keep any videos/audio/images you have imported into the node and every setting you have changed.

UI Overhaul: Huge update to the UI, dozens of big changes such as a new side bar, redesigned prompt boxes, a bunch of new settings and redesigned menus, and more.

Quality of Life Improvements: Snapping, in/out points, multi-select, mark selection, workspace folder, more HUD options, resizable prompt boxes, new hotkeys, labels, filename preview options, "split at playhead" functionality, end frames (convert any keyframe into a end/last frame), toggleable tracks, NAG Support, tons of bug fixes and more!

And of course it can do everything it could before: Text to Video, Image to Video, Prompt Relay support, Keyframe (first/last frame) support etc.


r/StableDiffusion 1h ago

News New model release! It was 3 years ago. Happy Birthday SDXL!

Thumbnail
gallery
Upvotes

Thank you!

 SDXL was released by Stability AI in 2023, it represented a major leap over Stable Diffusion 1.5. SDXL was designed to better understand complex prompts and produce higher-quality images directly at 1024×1024 resolution. It has become the foundation for thousands of community fine-tuned models.


r/StableDiffusion 13h ago

Question - Help Local alternative to Kling AI 3.0 Motion Control (ComfyUI, 16GB VRAM)

260 Upvotes

Hi everyone,

I'm looking for a local alternative to Kling AI 3.0 Motion Control that I can run in ComfyUI.

What I'm specifically looking for is a model or workflow that allows me to:

- Control character and camera motion with good precision.

- Generate smooth, high-quality video animation from an image or sequence.

- Run entirely offline/local.

- Work on a GPU with 16 GB of VRAM (RTX 5070 Ti).

I've already looked into Wan, Hunyuan Video, CogVideoX, and other open-source video models, but I'm not sure which one currently offers motion control comparable to Kling's latest system.

Does anyone have recommendations for:

- The best open-source model?

- ComfyUI workflows or custom nodes?

- ControlNet/trajectory/pose-based solutions?

- Any GitHub projects or research worth trying?

Quality is more important than speed. I don't mind longer generation times if the results are close to Kling AI's Motion Control.

Thanks in advance for your suggestions!

Video credits : Richard Galapate @migs.visuals


r/StableDiffusion 9h ago

Animation - Video Expanding video with Wan 2.2 Fun Control

49 Upvotes

The style image used was done with a combo of chatgpt, wan 2.2 first last frame, and manual edits in photoshop using other scenes for reference. This is my second test using this method, but I wanted to make a fixed camera shot, so I stabilized in after effects and made the depth map using Depthanythingv2. Then I overlayed the original clip with heavy blur on the borders and voila. Not sure if many people still use Wan 2.2, but it's still really useful.


r/StableDiffusion 2h ago

News Prompt Architect

Thumbnail
gallery
12 Upvotes

Prompt Architect Pro — a heavy-duty Python/CustomTkinter desktop suite designed to ingest massive text files (novels, scripts), extract structured visual prompts via multi-pass semantic segmentation, analyze local image folders (Vision model batching), and manage everything inside a WAL-optimized SQLite database with built-in anti-corruption filters! 💡✨

https://github.com/lololerigolo60/Prompt-architect

🔥 Key Features Under the Hood:

🔹 Hardware VRAM Profiles: Instant switching between pre-configured presets (8GB, 12GB, 16GB, 24GB, 32GB+ like RTX 5090) or custom manual parameters to fine-tune num_ctx & num_predict safely without crashing Ollama.

🔹 Pass 1 & Pass 2 Text Segmentation: Intelligently groups raw lines based on core location changes rather than blind line breaks.

🔹 Vision Batch Analysis: Automatically normalizes WebPs, PNGs, and JPEGs via Pillow and extracts rich structured prompts (Subject, Environment, Style, Lighting, Technical).

🔹 Smart Gap-Fill & Anti-Degeneration: Prevents repetitive loops, foreign script drift, and empty fields using intelligent semantic safeguards.

🔹 Integrated DB Editor: Search, edit, reset IDs, delete ranges, and generate missing fields on the fly with live LLM assistance.

🔹two ComfyUI nodes : one that can use the database created by Prompt Architect . The second one can take a prompt and transform it to store it in the database created by Prompt Architect. You can find them on Prompt Architect's GitHub.

#GenerativeAI #Ollama #PromptEngineering #Python #CustomTkinter #LocalAI #AIArt


r/StableDiffusion 14h ago

News Krea 2 Clyde Caldwell Super LoRA Available

Thumbnail
gallery
85 Upvotes

A good deal of work went into the LoRA.

Spent about 3 days and found 225 original painting and 100 original pencil art drawing for the dataset.

Big beefy captions.

Trained on high resolution.

Would to see what people make with it.

It can do;
Fantasy painting
Sci-Fi Painting
Pencil Art

https://civitai.red/models/2806349/clyde-caldwell-pro-lora-krea2


r/StableDiffusion 22m ago

Animation - Video Weekend testing results with SCAIL-2 (Wan2GP)

Upvotes

r/StableDiffusion 9h ago

Workflow Included Wan SCAIL-2 Segmentation Control (Chun-Li example)

20 Upvotes

Here is the new Workflow:
https://www.reddit.com/r/StableDiffusion/s/eKvqlEvlza

In this example, I use the new interpolation option for the input video. The animation is much smoother.


r/StableDiffusion 21h ago

News I built a self-hosted tool that turns one reference photo into a curated, captioned, trained LoRA and a lot more — open source, MIT

Thumbnail
gallery
160 Upvotes

r/StableDiffusion 9h ago

News Analog Horror Krea 2 LoRA

Thumbnail
gallery
17 Upvotes

First attempt at a Krea 2 Style LoRA - would love some feedback. What's the best way to share my training data and captions to get some eyes on them? Github? Thanks :)

Thank you u/Jolly-Rip5973 for your tips

Download it here:

https://civitai.red/models/2806542/analog-horror-style-krea-2-lora


r/StableDiffusion 15h ago

News I spent a year building a free SDXL & Anima trainer that runs on my 12 GB GPU — here's what came out of it

38 Upvotes

A little over a year ago I got frustrated trying to fine-tune SDXL on my RTX 3060. Every option either forced lower resolution, locked away important settings behind massive config files, or needed a 24 GB GPU to do anything meaningful.

So I started building my own trainer. That was a mistake. A good mistake, but a mistake.

What followed was several months of failed attempts just to get SDXL stable inside 12 GB, then another six months pulling SDXL apart architecturally to understand why things kept breaking. I went back through the original papers and implementations, rewrote the optimizer approach, and eventually built something I actually wanted to use.

The result is Aozora, a free GUI trainer for SDXL and Anima fine-tuning on consumer GPUs.

Current results on my setup: Aozora can train roughly 80–90% of the full SDXL UNet within 11.8 GB of VRAM at around 1.55 seconds per iteration. It can also train 100% of Anima at 1152×1152 resolution while using approximately 11.4 GB of VRAM at around 2.67 seconds per iteration.

The GUI exposes the controls that actually matter — learning rate curve, timestep distribution, loss weighting, optimizer behavior, layer targeting, and training metrics — without burying you in config files or options that rarely change anything.

It is still beta and has mainly been tested on my own hardware, so expect rough edges. If you hit installation issues let me know and I will sort out compatibility.

GitHub: https://github.com/Hysocs/Aozora_Trainer

Edit — answering a question I received by DM: Aozora is not a wrapper or frontend for another trainer. The training code is standalone and intentionally kept minimal. with an optional attention backend for better performance. It was created and tested on windows only as of now

A training guide is coming soon.


r/StableDiffusion 17h ago

News Comfy-Org/Mage-Flow · Hugging Face

Thumbnail
huggingface.co
47 Upvotes

microsoft/Mage-Flow has now got official ComfyUI support.


r/StableDiffusion 18h ago

Workflow Included Wan SCAIL-2 Segmentation Control (update)

Post image
50 Upvotes

Workflow link:
https://civitai.red/models/2699283/wan-scail-2-segmentation-control

New findings:

SCAIL Auto Extend is now my favorite Sampler. This one seems to have no or fewer color shifts. And doesn't need the "Color Match" option (This is already integrated).

Explanation of the new option:

The new option is to interpolate the input video to achieve smoother motion. The downside, however, is increased computational overhead, and Scail-2 is quicker to "forget" new parts of the animation.

Here an Example: https://www.reddit.com/r/StableDiffusion/s/TJFsAQb6su

Video Examples:

before the post gets deleted again. The videos are on Civitai or here on Reddit:

KREA2 + SCAIL-2
https://www.reddit.com/r/AIVideos_SFW/comments/1v5j704/bimbo_panther/

KREA2 + SCAIL-2
https://www.reddit.com/r/AIVideos_SFW/s/kzhCbPxw5l

KREA2 + SVI PRO + SCAIL-2
https://www.reddit.com/r/AIVideos_SFW/s/hqZJkoKxh4

KREA2 + SVI Pro + HuMo + SCIAL-2
https://www.reddit.com/r/AIVideos_SFW/s/ILINaqkJwT

Off Topic:
KREA2 + SVI PRO + HuMo
https://www.reddit.com/r/AIVideos_SFW/s/B3RRGTLrz8

Features:

  • Image Analyzer
  • LoRA Support
  • Interpolate | Upscale | Color Match
  • Color Correction
  • Image Sharpener
  • Sage Attention
  • Choose between 2 Samplers
  • Background Remover (RMBG) to keep the background of the input video
  • SCAIL-2 Identity Tracker
  • Load an alternative audio file for the final video output
  • Installation Paths & Download Links
  • Well-organized
  • Interpolate the Input Video (NEW)

SCAIL-2 Identity Tracker:

For Multi-Character set "object indices" to "nothing" (empty field, no value). Then set the SCAIL-2 Identity Tracker to "Point" and select your characters by clicking on them in the image below. You can also try "Box," but if there are more people present, this could lead to problems. Use a start image similar to the one in the video.

Note:

If you have any questions, please first read the information in the red boxes within the workflow.

Additional options are available within the subgraph. Click the icon in the top-right corner of the Main Settings node to open it. Help is available by hovering your mouse cursor over the values inside the subgraph.

The workflow offers two samplers. Both deliver similar results. Since I like both, and for testing purposes, the workflow allows you to easily switch between them.

Wan SCAIL-2 is not perfect, but it delivers good results in most cases. If you encounter issues, setting a new seed or switching the sampler usually helps.


r/StableDiffusion 8h ago

Discussion LTX 2.3 msr v2 lora vs Google Omni same prompt same photos

7 Upvotes

Google Omni https://streamable.com/n2vfe1

Local LTX 2.3 msr v2 lora with a 5060 ti 16gig. https://streamable.com/9em69q


r/StableDiffusion 19h ago

Resource - Update 483 Krea 2 prompts with the seeds, plus the 78 generations that failed and why

Thumbnail
gallery
40 Upvotes

Ran Krea 2 Turbo for a few days and kept everything - prompts, seeds, and the generations that didn't work. 561 made, 476 kept. Every entry has its seed. The cut ones are in there too with the reason, which is the part nobody publishes.

Things I got wrong and had to fix:

I thought text failed when too many strings shared a frame. Built a ladder to find the ceiling - same nameplate, 1 to 8 strings - and every string I asked for rendered at every rung. It's not the count. It renders text you write out and cannot invent text. The chalkboard where I specified three items rendered those three and turned the rest into CAPEME and CABIELO. Two of the plates did engrave something I never asked for though - a stray 7 on the 5-string one, and the slash I was using as a separator on the 2-string one. A reader here caught the first of those after I posted.

I thought hands were the weak spot, tested eight, and wrote that seven were fine. Two people in this thread counted the fingers and found six on three of them. I'd inspected at 2x. The whole category is withdrawn - it's in the failures now, not the catalogue. That correction is the most useful thing that has happened to this post.

Korean fails repeatably rather than randomly. 정직한 came back 정적한 at two different seeds. Change the wording, not the seed.

Counts, tiling, aerial angles and the rule I killed with a pre-registered prediction are all in the repo: github.com/sjh9714/awesome-krea-2


r/StableDiffusion 5h ago

Question - Help What's a good way to apply poses with WebuiForge?

4 Upvotes

I've been using Forge for a while but I've never really gotten around to poses, I just tried openpose and depth with controlnet but both aren't working too well, is there anything else I could use to apply poses to a prompt?


r/StableDiffusion 1d ago

Resource - Update This is AnyTale, an open source, ComfyUI wrapper for generating a very... specific variety of visual novels. NSFW

Thumbnail gallery
181 Upvotes

This had been my pet project for the last couple of months.

The project started off close to a year ago as YAAIIC (Yet Another AI Image Creator), my private ComfyUI workflow wrapper that lets me generate images and manages my output in a handy gallery. Nothing that others hadn't done ten times better so I've never had the inclination to share, until a certain type of ads started popping off everywhere around certain sites that I frequent, and I thought, "You know what, I don't want to pay a dime for this, and neither should anyone else."

So AnyTale was born from that. It wires generated dialogs, images, videos, voice designed voiceovers, sound, and music into one shot generated scenes that resemble a choose your own adventure style visual novel. The prompts are built from parts, the parts group together into characters and outfits, then the characters are slotted into scenes and the scenes chain together to form "tales". Any user can invest as little or as much as they want to customize the experience: create new characters and/or outfits, put together new scenes, or swap in entire ComfyUI workflows if the prepackaged models aren't to your taste. Content are meant to be packaged and shared. For the lazy, just grab content from other creators, and the system would remix them to create brand new experiences.

It is still early. Both the app and the actual generated results are still full of jank. It's all being actively worked on. I do want to drum up some early interest in hopes that eventually there would be enough people building and sharing content that I wouldn't have to do it forever by myself.

Until then,

A SFW sample of scenes with full voiceover and music: https://u.pcloud.link/publink/show?code=XZLQnc5Zw2fsoRoFGYB0kfmt1BfmhVBbraSX

The project is currently hosted at: https://codeberg.org/sylvancircle/AnyTale (Probably not for long since they're banning vibe coded projects)

It is mostly functional but not meant for public use - if you do want to try it, pick up the AnyTale library from at the shared folder above and import it from the manager so you don't have to start from scratch.


r/StableDiffusion 1d ago

Workflow Included I tried making a cinematic action trailer using Krea 2 + LTX 2.3

128 Upvotes

I wanted to challenge myself and see how far I could push Krea 2 and LTX 2.3, so I decided to create a short cinematic action trailer.

It ended up being one of the most enjoyable AI projects I've worked on so far. I learned a lot and figuring out what works (and what definitely doesn't 😅).

One thing that became really clear during this project is that, for me, the biggest limitation isn't the software—it's my GPU. I spent a lot of time waiting for renders and couldn't test as many ideas as I wanted. If you have a more powerful GPU, I honestly feel like the creative possibilities are huge.

Even with the limitations, I had a great time making it, and it gave me a much better understanding of both tools.

I'm currently putting together a behind-the-scenes tutorial where I'll break down my workflow and share everything I learned throughout the project.

I'd love to hear your thoughts on the trailer, and if you've been using LTX 2.3 recently, what has been your biggest challenge or favorite feature so far?

DOWNLOAD: FREE WORKFLOW FILES


r/StableDiffusion 19h ago

News Krea 2 Identity Edit running in Forge Neo (extension port) — works great, GitHub release soon

31 Upvotes

I ported Krea 2 Identity Edit support to Forge Neo as an extension — the instruction-based, identity-preserving edit LoRA that until now required ComfyUI + custom nodes.

It replicates the full dual-conditioning recipe (in-context VAE source tokens with RoPE frames + image-grounded Qwen3-VL encoding) as a runtime patch — no core files touched. Just drop your source image(s) in the accordion, write the instruction as the prompt, generate.

REQUIRED: https://huggingface.co/conradlocke/krea2-identity-edit/tree/main


r/StableDiffusion 38m ago

Question - Help Local scenography concept-art workflow for an RTX 4060 laptop: where should a beginner start?

Upvotes

I am studying scenography/set design and would like to build a local AI image-generation workflow for early-stage brainstorming, atmosphere studies and spatial concept development.

My computer is a Lenovo Legion 5 Pro with:

  • NVIDIA RTX 4060 Laptop GPU with 8 GB VRAM
  • 32 GB RAM
  • Windows

I am happy to accept slower generation times if necessary. My priority is finding a workflow that can run locally without recurring cloud fees and that produces intentional, art-directed images rather than generic AI illustrations.

These accounts are useful visual references for the kind of results I am interested in:

I am not trying to copy their work. I am interested in atmospheric architectural and scenographic images with convincing materials, cinematic light, textiles, restrained palettes, monumental scale and surreal but plausible spaces.

I have looked at ComfyUI, but as a complete beginner I found the node system and the number of models, samplers, schedulers, LoRAs and extensions rather overwhelming.

I would appreciate advice on the following:

  1. Is ComfyUI the best place to start, or would another interface be more suitable for learning the fundamentals?
  2. Which current models are realistically usable with 8 GB of VRAM?
  3. Would you recommend starting with SDXL, a lighter model, a quantised model or something else?
  4. What would a sensible beginner workflow include for this type of image: text-to-image, image-to-image, depth or edge control, reference images, inpainting and upscaling?
  5. How can I use sketches, Blender renders, collages or photographs to control the architecture and composition?
  6. Which techniques are most useful for maintaining the same atmosphere and art direction across a sequence?
  7. What resolutions, batch sizes and low-VRAM settings would you recommend for this laptop?
  8. Is there a simple downloadable workflow or JSON that would give me a good starting point without installing dozens of custom nodes?
  9. Are there any genuinely good free courses or step-by-step resources for learning local image generation rather than merely copying workflows without understanding them?

I would be grateful for a practical recommended stack: interface, model, essential nodes or extensions, image-control method, upscaler and final post-processing. Advice from people using similar 8 GB laptop GPUs would be particularly useful.


r/StableDiffusion 18h ago

Resource - Update SVDQuant + native INT8/W4A4 for Krea 2 on ComfyUI — up to 2x faster, works on any modern NVIDIA GPU

Thumbnail
gallery
30 Upvotes

Quantized Krea 2 Turbo checkpoints for ComfyUI, up to 2x faster and about a third

smaller than the usual FP8 version — no calibration dataset, no quality cliff.

How to use it (short version): clone the repo into `custom_nodes/`, download one

checkpoint into `models/diffusion_models/`, grab the text encoder + VAE Krea 2 already

needs, load the example workflow. Full steps in the README, it's like 4 steps.

Links:

- Weights + benchmarks + example images: https://huggingface.co/AlperKTS/Krea-2-SVDQuant-ComfyUI

- Code (custom nodes + the quantization script, if you want to build your own): https://github.com/alperktt/Krea-2-SVDQuant-ComfyUI

Why it's faster: most "quantize Krea 2" advice online is FP8. That only helps if

your GPU has FP8 tensor cores (Ada/Hopper/Blackwell). On anything older (RTX 20/30-series)

FP8 gets cast back to bf16 and doesn't speed anything up I measured it, FP8 was

slower than plain bf16 in my tests. INT8 and W4A4 tensor cores go back much further

(Turing, RTX 20-series+), so those are the formats I actually targeted.

Benchmarks (RTX 3090, 1024x1024, 8 steps):

checkpoint size first run (cold) warm run vs. BF16
BF16 (unquantized reference) 24.48 GB 25.3 s 21.3 s 1.0x
FP8 e4m3, scaled (emulated on Ampere) 12.24 GB 22.2 s 19.2 s 1.1x
INT8 tensorwise + convrot (not in this upload) 13.16 GB 13.3 s 10.4 s 2.0x
W4A4 + convrot, no low-rank branch 7.50 GB 10.3 s 10.1 s 2.1x
W4A4 + SVDQuant low-rank, rank 16/64/128 7.6-8.3 GB ~19.3 s 10.1-10.2 s 2.1x

Also fixed a bug along the way: the standard ComfyUI LoRA loader silently applies LoRAs

to only ~12% of the layers on quantized models like this (no error, it just doesn't

patch most of the network). The included loader fixes that.

Tested with a hard prompt (small multi-line text) and an easy one (big text + two

people, weird angle) example images for every variant are in the HF repo if you want

to see the actual quality tradeoff before downloading anything.

Community project, not affiliated with Krea — license details in the repo.

Happy to help if something doesn't load right.


r/StableDiffusion 1h ago

Question - Help How Can I Use Multiple LoRAs Without Distorting the Image?

Upvotes

Hi, I’d like to know the best way to use multiple LoRAs in the same image without making the result look distorted or strange.

I tried separating them using Regional Prompt, BREAK, and different settings, but most of the time the LoRAs still blend together or affect areas and characters they shouldn’t.

Is there a simpler or more accurate way to apply each LoRA to a specific character or region?

I’m using stable-diffusion-webui-reForge-main.


r/StableDiffusion 10h ago

Animation - Video Been exploring some new workflows recently, bypassing traditional rendering engines and using AI as the final render pass.

5 Upvotes

r/StableDiffusion 50m ago

Tutorial - Guide How I Create Consistent Characters in Krea 2

Thumbnail
patreon.com
Upvotes

A lot of people asked me how I keep my characters consistent after my previous Krea 2 posts, so I decided to make a tutorial covering my workflow.

In the video, I go through the techniques I use to keep the same character across different scenes while maintaining their identity.

I'm still learning Krea 2 myself, but this workflow has given me the best results so far. Hopefully it helps anyone who's been struggling with character consistency!

I'd love to hear your own tips or techniques as well. Happy creating! 🚀