#VideoStages - Take a video from first draft to finished clip in a few simple stages

1 messages · Page 1 of 1 (latest)

rocky kite
#

VideoStages adds a multi-step video flow to SwarmUI. Instead of asking one generation to do everything at once, you can build on your result stage by stage to improve motion, detail, and overall polish while keeping the whole process in one place. If your workflow also creates audio, VideoStages automatically carries that into the finished video too, including audio from AceStep and, soon, Qwen-TTS.

Think of it as draft, refine, and polish for video, built right into the normal SwarmUI experience.

Works well with TTS or Music models. See my other extension AceStepFun to get started genning your own 1girl music videos.

If using LTX2.3, here's some resolutions I've found that work well for using the 1.5x upscaler, which requires dimensions always be divisible by 16 (height and width, before and after upscaling):

  • 768×1024 -> 75% -> 192×256 -> upscale 1.5x -> 288×384

  • 768×1024 -> 75% -> 192×256 -> upscale 1.5x -> 1.5x -> 432×576

  • 768×1024 -> 75% -> 192×256 -> upscale 1.5x -> 2x -> 576×768

  • 768×1024 -> 75% -> 192×256 -> upscale 1.5x -> 1.5x -> 2x -> 864×1152

  • 768×1024 -> 50% -> 384×512 -> upscale 1.5x -> 576×768

  • 768×1024 -> 50% -> 384×512 -> upscale 1.5x -> 1.5x -> 864×1152

  • 768×1024 -> 50% -> 384×512 -> upscale 1.5x -> 2x -> 1152×1536

  • 768×1024 -> 50% -> 384×512 -> upscale 1.5x -> 1.5x -> 2x -> 1728×2304

  • 768×1280 -> 50% -> 384×640 -> upscale 1.5x -> 576×960

  • 768×1280 -> 50% -> 384×640 -> upscale 1.5x -> 1.5x -> 864×1440

  • 768×1280 -> 50% -> 384×640 -> upscale 1.5x -> 2x -> 1152×1920

  • 768×1280 -> 50% -> 384×640 -> upscale 1.5x -> 1.5x -> 2x -> 1728×2880

  • 1024×1536 -> 75% -> 256×384 -> upscale 1.5x -> 384×576

  • 1024×1536 -> 75% -> 256×384 -> upscale 1.5x -> 1.5x -> 576×864

  • 1024×1536 -> 75% -> 256×384 -> upscale 1.5x -> 2x -> 768×1152

  • 1024×1536 -> 75% -> 256×384 -> upscale 1.5x -> 1.5x -> 2x -> 1152×1728

  • 1024×1536 -> 50% -> 512×768 -> upscale 1.5x -> 768×1152

  • 1024×1536 -> 50% -> 512×768 -> upscale 1.5x -> 1.5x -> 1152×1728

  • 1024×1536 -> 50% -> 512×768 -> upscale 1.5x -> 2x -> 1536×2304

  • 1024×1536 -> 50% -> 512×768 -> upscale 1.5x -> 1.5x -> 2x -> 2304×3456

delicate crypt
#

Gave it a try and it is nice. Tho sometimes it seamed to add uncececery VAE decode into VAE encoder in between LTX2 stages where latents could just flow on.
What is the recomended settings setup to reproduce the official LTX2 workflow? (generate at 0.5x then 2x latent upscale then 2nd stage generation at full res for final output)

rocky kite
#

VideoStages - Take a video from first draft to finished clip in a few simple stages

delicate crypt
#

But what is also the full setup? Should it be ImageToVideo and then add a stage using video stages? Should a Image be input and 2 video stages in the plugin? What should the Control be set for them? Is the latent upscale supposed to be in the first stage or second stage?...etc

rocky kite
#

Whatever you want

#

Are you asking for what I use?

#

That'll gen your tiny preview video. If you decide to upscale it, just click these two

#

In the above examples, Video Frames is unset, because I'm also using my AceStepFun extension to create music using AceStep, so the "Connect Audio to Video" option in VideoStages automatically detects the audio track and calculates the correct video length (total frames)

delicate crypt
#

Thanks, was just wondering what is the correct way to use the extension, so that i don't use it wrong and then complain about it being bad

icy sonnet
#

I don't see this in the Extensions list - how do I install?

rocky kite
icy sonnet
#

ah gotcha will do tnx

rocky kite
#

Here's a few examples of how starting a tiny 256×384 video and upscaling twice at 1.5x creates a coherent, higher quality 576×864 video

The initial video takes ~11 seconds, the upscaled video takes ~60 seconds

The third video is upscaled at 1.5x three times, takes 171 seconds to generate and results in a high quality 864×1312 video - but you'll notice it has a white border at the bottom because the ltx2.3 1.5x upscaler starts getting bothered with the resolution sizes.

Doing 1.5x -> 1.5x -> 2x avoids the border but takes a bit longer than 171 seconds

rocky kite
#

Made some changes.

  • VAE Decode Tiled params can now be set via the Advanced Sampling menu
  • LTXVImgToVideoInplace.strength value can now be set using <param[Video Stages LTXVImgToVideoInplaceStrength]:1>
#

I guess I should work on wan support next. so boring. who even uses wan

rocky kite
#

I see that comfyui has an official workflow for first frame/last frame with ltx2.3

Looking at it now to see if it makes sense to implement in this extension

delicate crypt
#

Multiframe also works, wouldn't mind seeing an extension for it

rocky kite
#

I am glad to be contributing to the downfall of society, in my own little way

delicate crypt
#

People are already making heaps of money by posing as hot girls on onlyfans using AI generated photos

icy sonnet
#

That extra hand must be handy.

#

When adding a preset the VideoStages options don't all appear

rocky kite
icy sonnet
#

gotcha.

rocky kite
#

Here's a stronger list of starting resolutions that are compatible with the 1.5x upscaler:

  • 768×1024 -> 75% -> 192×256 -> upscale 1.5x -> 288×384

  • 768×1024 -> 75% -> 192×256 -> upscale 1.5x -> 1.5x -> 432×576

  • 768×1024 -> 75% -> 192×256 -> upscale 1.5x -> 2x -> 576×768

  • 768×1024 -> 75% -> 192×256 -> upscale 1.5x -> 1.5x -> 2x -> 864×1152

  • 768×1024 -> 50% -> 384×512 -> upscale 1.5x -> 576×768

  • 768×1024 -> 50% -> 384×512 -> upscale 1.5x -> 1.5x -> 864×1152

  • 768×1024 -> 50% -> 384×512 -> upscale 1.5x -> 2x -> 1152×1536

  • 768×1024 -> 50% -> 384×512 -> upscale 1.5x -> 1.5x -> 2x -> 1728×2304

  • 768×1280 -> 50% -> 384×640 -> upscale 1.5x -> 576×960

  • 768×1280 -> 50% -> 384×640 -> upscale 1.5x -> 1.5x -> 864×1440

  • 768×1280 -> 50% -> 384×640 -> upscale 1.5x -> 2x -> 1152×1920

  • 768×1280 -> 50% -> 384×640 -> upscale 1.5x -> 1.5x -> 2x -> 1728×2880

  • 1024×1536 -> 75% -> 256×384 -> upscale 1.5x -> 384×576

  • 1024×1536 -> 75% -> 256×384 -> upscale 1.5x -> 1.5x -> 576×864

  • 1024×1536 -> 75% -> 256×384 -> upscale 1.5x -> 2x -> 768×1152

  • 1024×1536 -> 75% -> 256×384 -> upscale 1.5x -> 1.5x -> 2x -> 1152×1728

  • 1024×1536 -> 50% -> 512×768 -> upscale 1.5x -> 768×1152

  • 1024×1536 -> 50% -> 512×768 -> upscale 1.5x -> 1.5x -> 1152×1728

  • 1024×1536 -> 50% -> 512×768 -> upscale 1.5x -> 2x -> 1536×2304

  • 1024×1536 -> 50% -> 512×768 -> upscale 1.5x -> 1.5x -> 2x -> 2304×3456

crystal cedar
#

This looks cool! Does it also let you merge videos together?

rocky kite
#

no

crystal cedar
#

Awe dang, is that just not a thing people tend to care about?

rocky kite
#

I just haven't done it

#

when you say merge ... ?

crystal cedar
#

Ive been looking for something like the svi pro 2 sort of deal that uses last image to string together videos with ffmpeg. I remember it had some sort of image fix process in between to counteract wans degradation so i thought maybe this might be able to do something similar

rocky kite
#

just provide me some non-insane workflows to look at and I can def. add it in

crystal cedar
#

That'd be awesome, i'll see if i can find the workflow again, i never ended up downloading it before cause comfy scares me too much 😄

crystal cedar
delicate crypt
#

The WAN SVI was never officialy supported in SwarmUI, but it does work really well. It lets you do 1 minute long Wan videos without any degradation issues. Now there is much less need for that because LTX2 doesn't suffer from the same degradation problem. (tho admitedly for some specific things Wan still works better)

rocky kite
#

well cus it looks like it's all custom stuff

delicate crypt
#

Yeah StableVideoInfinity needs some latent space dickery to insert the reference frames, so it doesn't work with just official Comfy nodes. But later on the actual sampling is standard Comfy Wan stuff

rocky kite
#

well it definitely fits within the spirit of VideoStages' goal, so if you create a sample workflow I can follow along with I can add it in

#

using swarmui and native comfyui nodes where appropriate, I mean

delicate crypt
#

Il have a look trough my Wan SVI workflows to see if i got a good simple example. Usually they look pretty tangled because they are a chain of video extend workflows

rocky kite
#

what a surprise

delicate crypt
#

Tho i never looked into what that node actually does. It might be possible to reimplement that node out of a tangle of more standard nodes

delicate crypt
#

This is what this node does

#

I don't touch comfy guts to know exactly what it does, but from what i guess it inserts an extra frame containing the anchor image (that is masked so it can't be changed) while also copying over a few frames worth of motion latent data from the end of the previous video extend segment

crystal cedar
#

Oh wow that looked like a nightmare in the first one lol. I have no idea how it works like as far as comfy goes but from what i read it uses some sort of math pattern that counteracts the degradation by gradually shifting things in the opposite direction at the same rate the degradation would shift it normally. If that makes sense, it only barely made sense to me and i might not be awake enough today to explain it properly lol

delicate crypt
#

My understanding is that they give it an example image of "Here is what the video should look like" and then trained a LoRA model to look at that specific anchoring frame and tries to modify the rest of the generated frames to look more like the anchor frame. So that as Wan is gradualy destroying the image quality, the LoRA is subtly repairing it slightly, just enugh to prevent Wan from spiraling down into a incoherant mess(as it normally does).

#

And by the look of it i have fed it the first frame as the anchor frame, not the last frame of the previus video segment. This might be a mistake, but the workflow still worked fine (because the character always stayed in frame)

delicate crypt
#

So it might be possible to prepare the correct latent space using more standard comfy nodes (at least with an uncececery VAE decode then encoder), but the problem is that would likely not carry over the motion latents. So you would keep the video from degrading, but you would still get the motion discontinuity between clips as in the bad way of video extend where the last frame is simplfy fed on as the first frame. (Proper SVI workflow does not suffer from motion discontinuity problems)

crystal cedar
#

Ahh okay, that sounds like a much better way to explain it! In my gens currently i still get the distortion over time on top of the motion being off during transitions(sometimes messing with the prompt or timing can help avoid that like inserting a 2 sec video for a smoother transition before going back to 5 sec) . Its gotten a lot better with using more recent models for sure, and with tweaking stuff, but after a few continuations the degradation always still adds up pretty bad beyond a certain point. It's kinda like it keeps degrading faster each time after that so if i try to like get to a minute long its always very blurred by then.

delicate crypt
#

Then you perhaps do want to do it how i accidentally did it by using the first frame of the first clip as the anchor latent (not the last frame of the previous clip)

rocky kite
#

I'm hesitant to bring in kijai node deps just for that one node

#

but it looks like it's ... just conditioning the input image before handing off to the ksampler?

#

obviously with latent motion and some other stuff. funny this isn't part of comfyui core, as the ltx stuff is

delicate crypt
#

I think you just need to insert the anchoring image as the first frame of the video and mask it so it can't change, then give it the actual first frame as the 2nd frame in the video. That should make it work for img2video i think

#

But for video extend workflows we would need to copy over some motion latents, and i don't think that is possible to do with any standard comfy nodes

rocky kite
#

no way that stupid file is 3,300 lines long

delicate crypt
#

Well it has nodes in there for running like 10 different video models. Kijai just throws in there any new video model that gets released with minimal porting

rocky kite
#

yes I'm well aware of how sloppy this thing is, but it brings joy to many a gooner so who am I to complain

delicate crypt
#

Unless you want to create your own comfy custom node pack and just steal the relevant nodes

rocky kite
#

ok so the length is split into 5 seconds maximum (81 frames) per stage, so we can automate that

#

😂 I don't know if steal is the right word to use

delicate crypt
#

Then il call it exercising your rights provided by the projects Apache License 2.0 😄

rocky kite
#

man I have not been able to get SVI to work

#

this is using the KJ nodes, which I finally gave in and installed, just to see what I was messing up

#

because I had set it up to not use any custom nodes and I was still getting the above - I thought I had some bug I couldn't find. After spending ~3 hours of the damned thing I gave up yesterday

#

but ... KJ does the same thing

rocky kite
#

I'm marking this one as won't do

crystal cedar
#

Dang, that's cool you gave it a solid shot though! I wonder what makes it so particular.

delicate crypt
#

Yeah it do be like that sometimes.
Last time i couldn't get LTX2.3 frames to video working (made noise or a horror show). Then got a workflow for it and it did the same thing yet it worked somehow.

#

Btw do we have any good LTX2.3 video extend options in SwarmUI yet?

stuck rampart
#

I like that VideoStages upscaling works for Wan outputs too, but I've encountered a small hiccup: for Wan videos, when I enable frame interpolation under Advanced Video, it seems to enforce 24 FPS on the VideoStages output. Because my base Wan video is 81 frames at 16 FPS, this shortens the VideoStages output from 5 seconds to about 3. 🤔

In order, the videos are:

  1. Base Wan output - 0:05 duration
  2. Wan with VideoStages enabled (1.5x upscale) - 0:05 duration
  3. Wan with VideoStages (1.5x upscale) and 2x frame interpolation - 0:03 duration
pallid aurora
#

Restarted swarm this afternoon and video stages is failing to build.

rocky kite
#

gimme gimme gimme da error

pallid aurora
#

no accessible extension method ‘AltLatent’ accepting a first argument of type ‘WorkflowGenerator.ImageToVideoGenInfo’ could be found.

rocky kite
#

I might need a few hours before I can look into this

pallid aurora
#

Earlier: ‘WorkflowGenerator.ImageToVideoGenInfo’ does not contain a definition for ‘AltLatent’

#

At first glance I’m like who is inside who here? lol.

#

Yup, just wanted to make sure you knew.

#

Thanks.

icy sonnet
rocky kite
#

hooray for a strong unit test suite

#

@pallid aurora @icy sonnet fixed

hot garnet
#

Okay dang I need like... a GitHub Action that pulls and compiles all registered extensions after swarm commits to catch these quicker

rocky kite
hot garnet
#

-# i hate github action development cause you need to push a commit to github to test each change lol

rocky kite
#

and in truth, an auto-pull and build should be part of each separate extension

icy sonnet
#

Thanks guys.

crystal cedar
#

after uninstalling and trying to reinstall, it seems like videostages is still there but not on the list so the install button doesn't work. i also get an error on startup now regarding videostages

icy sonnet
#

Can confirm same issue. I don't see VideoStages in the extension list anymore, when I hit install it says it's already installed.

#

Let me see if I can manually re-move it and reinstall it

#

@crystal cedar shutting down, removing the extension directory from SwarmUI\src\Extensions, then installing again, (refresh browser) worked

#

The sub directory under there not the entire Extensions directory

crystal cedar
#

sweet, thanks! that solves the last of my recent combo of errors and/or personal blunders hehehe

rocky kite
#

I don't know how you'd get into that state, unfortunately

rocky kite
#

I'm working on VideoStages and Base2Edit so they can work a bit more closely together.

  • add width/height field to the root-level VideoStages section. If both are defined, the image is resized to those dimensions before the first native video stage
  • allow selecting where the init image comes from: base/refine/edit stage

From this, we'll be able to do a FFLF workflow!

rocky kite
#

I may not have thought this all the way through

crystal cedar
#

Dang that sucks you hit a speedbump, even if i dont understand that image 😄

I'm rooting for ya! Hopefully its still doable, a way to avoid the extending issues would be huge.

rocky kite
#

not a speedbump, just more changes than I had anticipated, but it's working well right now in my testing: #gens message

crystal cedar
#

Nice! I'll go check some out in a bit thats good news

pallid aurora
#

@rocky kite did full update to swarm/comfyui/extensions. Swarm says everything is up to date, but now VideoStages won't build.

#

Did maybe it do updates out of order and now that it doesn't build, it won't pull the update it needs to pull for VideoStages?

rocky kite
#

you gotta update it manually now I guess

pallid aurora
#

just git pull?

rocky kite
#

yeah and restart server

#

big enormous update coming btw

#

scene-based

pallid aurora
#

yeah, something is borked in the update everything and restart swarmui thing...maybe...

#

oooooh

#

yup git pull:

$ git pull
remote: Enumerating objects: 128, done.
remote: Counting objects: 100% (128/128), done.
remote: Compressing objects: 100% (22/22), done.
remote: Total 96 (delta 76), reused 94 (delta 74), pack-reused 0 (from 0)
Unpacking objects: 100% (96/96), 35.00 KiB | 95.00 KiB/s, done.
From https://github.com/jtreminio/SwarmUI-VideoStages
   33ce0d5..d75a249  main       -> origin/main
 * [new branch]      flf        -> origin/flf
Updating 33ce0d5..d75a249
Fast-forward
 .gitignore                |  2 ++
 AGENTS.md                 |  4 ++++
 src/LTX2/StageExecutor.cs |  9 ++++-----
 src/StageRunner.cs        | 49 +++++++++++++++++++++++++++++++++++++----------
 4 files changed, 49 insertions(+), 15 deletions(-)

user2@WINDOWS-DA4MTCI MINGW64 /g/___all_webuis/SwarmUI/src/Extensions/SwarmUI-VideoStages (main)
rocky kite
#

did it work? should just need a rebuild I think

pallid aurora
#

Yup, it rebuilt. Now I just gotta wait for the cursory boot scan of 80k+ models.

#

Innnteresting...just tried a git pull on Qwen-TTS and got a bunch of updates too...did it get fixed as well? OMG IT BUILT

rocky kite
#

no

#

unless they updated it upstream

pallid aurora
#

Ah, so builds but will most likely throw an error on execution.

rocky kite
#

you can downgrade its transformers deps manually if you really need it

pallid aurora
#

i don't, i can wait.

plucky wigeon
#

Whoa, this is complexicated. lol.

rocky kite
#

That's unfortunate, I tried to make it as easy as I could, more or less replicating existing swarmui process

plucky wigeon
#

I -think-I got it to work on my first try. Any chance of inheriting params from the basic gen? Like, steps, cfg, sampler, scheduler? Just an idea.

rocky kite
#

do you find yourself doing text-to-video or image-to-video more often?

plucky wigeon
#

Its definitely faster than generating it raw, did 480 with 2x upscale and it was 103.48 sec gen vs 3.52 min gen

plucky wigeon
#

I just wanted to test it so made rando girls talk about how they ended the clone wars, lol. Trying to stay up to date with this stuff.

rocky kite
#

This is what I'm working towards

edit: this redesign target has been cancelled.

plucky wigeon
#

Oh wow that's dope

rocky kite
#

the tricky part is splicing audio into chunks so for example the 3rd audio track there can be used in scene 1 and scene 2

pallid aurora
#

FYI, I mostly do T2I2V. I should probably figure out a better multi T2I, cherry pick smaller list for I2V.

#

But I like having all the metadata for repro in one place.

rocky kite
#

absolutely

#

@pallid aurora that is my most common workflow, and this extension is primarily aimed at me having fun with ai video generation, so I won't be doing anything overly complicated or advanced. The endgoal is and always has been to gen 1girl

delicate crypt
rocky kite
#

I apologize in advance to everyone who is using this extension, but the upcoming version will not be backwards compatible with the existing version's metadata.

#

Since I'm moving to a videoclip -> stages model, instead of stages attaching to a single root-level video clip, everything is an array now

#

I could write a migrator, I'll see if it's worth doing, I don't know how many people actually use this extension right now.

wary epoch
#

Bring it on 🙂

rocky kite
#

Might even release it today

#

Yes the UI is completely, utterly different to what I had been teasing, but the more I worked on it the less it made sense to go down that route, other than for overlapping audio clip support. I'll figure that one out later.

#

but, more crucially, it began diverging further and further away from native swarmui tools, and that's not something that is fun to manage

delicate crypt
#

I like the look of that UI, flexible but easy to understand

rocky kite
#

The redesign is now live.

pallid aurora
#

Whenever I reset parameters, sometimes just refreshing the page, the switch for the VideoStages parameter set auto-activates.

plucky wigeon
#

Same
Sometimes on a fresh launch

pallid aurora
#

tagging @rocky kite

rocky kite
#

Looking into this now, while I commit my first controlnet support into a branch

[controlnet e47e35e] Working, messy
 25 files changed, 1979 insertions(+), 38 deletions(-)```
😭
#

way too many changes, far more than I anticipated

#

@pallid aurora @plucky wigeon can you delete all the clips within videostages, and reload? I spent a lot of time trying to nail down all possible causes of that switch getting flipped, but it's difficult because it's dynamic json

plucky wigeon
#

All of the clips...

rocky kite
plucky wigeon
# rocky kite

Thats what I thought you meant but wasnt sure. I do gotta start messing with this.

pallid aurora
#

I delete them, then they come back.

#

It’s crazy making. I’ll manually hit generate a dozen times. No. Prob. Edit the prompt. Hit generate and suddenly that section is active again. Maybe it’s that whole video tag activates video, but videostages is catching it instead?

rocky kite
#

try to note the action that triggers it

#

it has a few watchers, like adding edit stages via base2edit, or acestepfun; I can move those to activate on dropdown click instead of at point of action

pallid aurora
rocky kite
#

I found the cause and fixed the issue. Unfortunately I did it on my controlnet branch, so it might not be until later tonight or tomorrow for when I merge it

misty stag
rocky kite
#

just seeing that message

#

now I'm hesitant to push this controlnet update

#

everyone thank PG

wary epoch
plucky wigeon
#

Not gonna lie, very confused about using this lol.

#

So like, for a basic usage.

#

You know how github pages have # Installationm # Usage. We need that.

#

This...

#

Sounds amazing, but where to even begin

#

Maybe have a few # Usage examples in 3 steps, one for upscaling and another for extending to get people started

delicate crypt
#

Yeah it would be great to have an setup example of Lightricks official LTX2 two stage workflow with the latent upsampler

rocky kite
#

Just pushed big update ... and I see this

#

I'll create some quick screenshots right now explaining LTX2.3 usage, then update README for later tonight

#

but, for the update I just pushed:

  • ControlNet support has been merged, for LTX2.3
  • first/last frame support has been added to WAN (LTX already supported). Only 0, 1, or 2 images can be used
plucky wigeon
#

Well, I wanted to try it out and give you some feedback but I honestly have no clue how to even start. >_<

#

I know you have been very verbose in discord about it as you worked on it, but I am attacking this from the angle of rando off the street

rocky kite
#

Here's the most basic LTX2.3 workflow you can do with the extension

This is a single video-clip, two-stage workflow with first frame and last frame.

plucky wigeon
#

So, do you use an init image with this setup, or text to everything?

rocky kite
#

either. You can add as many or as few reference images as you want, and set what frame in the video they are used for (for LTX).

For WAN you can either provide 0 images, first-frame image, or first and last-frame image

#

You can use global prompt, <video> or <videoclip> to define your prompt.

Further, you can use <videoclip[0]> to target the first video clip. It will be applied to all stages within that clip.

#

To clarify: you don't need to use Swarm's "Image to Video" section for this. You can if you want, and any settings you set there will be applied to VideoStage's config, but it's not necessary.

#

for controlnet support, add your controlnet source as normal

#

Choose your lora

#

your Video's stages will already have a new controlnet strength slider you can adjust

#

LTX has a ton of different options I needed to juggle and try to figure out how to best cram into the tiny sidebar. That's why I had originally began my redesign focused on the larger, wider bottom section

#

but, I feel that once you click around and gen your first video, it becomes a bit more obvious what's going on

#

Global: width, height, and FPS apply for everything.
Video Clips: discrete videos. What normally happens already in SwarmUI when you do an image-to-video or text-to-video workflow, it goes from A-Z and then saves a single video. What is different about VideoStages is you can have as many video clips as you want, and at the end it joins them together for one long video clip that contains multiple, separate videos.
Video Stages: Think of it like Base2Edit's stages. You're working on the same video clip, but upscaling it, refining it, etc, as many times as you want.

#

In the screenshots above, I've shown a single video clip, two stage workflow. Other than the controlnet, there's no difference between this example and what SwarmUI can already natively do.
What's different is you can upscale more than once with VideoStages, which is what I do all the time: begin with a tiny little video and if I like it I upscale it by 1.5x twice

#

With this week's redesign, you can now do the above, across any number of separate video clips.

That's how I've created this harlem shake video. There's a clear cut halfway through, that's the two video clips in action. Each video clip gets its part of the song

#

Yeah it's a stupid video but it was one of the first I tried

#

With today's controlnet update, you can feed it a source video, like this one

#

For now though, start with the most basic possible workflow:

  • one video clip, one stage, one init image (first frame)
plucky wigeon
#

What model chosen at the bottom, and any other video params chosen anywhere else?

rocky kite
#

I chose zit for my image model

#

I actually haven't tried text-to-video in some time! Let me try it now, I'm actually nervous

#

I've got an entire test suite, so in theory it should work as expected.

#

yeah, text-to-video works

#

Small bug: base/refiner options shouldn't be shown in a text-to-video workflow, only upload

#

but I'm tired boss. Just use image-to-video for now, which I would assume is the most common workflow

#

Oh, if you use Base2Edit, it auto-detects when you've added an image and it shows up in your Reference Image / Image Source dropdowns.

As well as AceStepFun appears in Audio Source dropdown

#

unfortunately you can't use swarm's native audio support right now because it actively blocks video gen in a text-to-audio workflow

plucky wigeon
#

Load the LoRa in the bottom?

rocky kite
#

yeah just like normal

plucky wigeon
#

Why is that red

#

We doing 0.6 distill?

rocky kite
#

yeah I've found that works best with ltx2.3

plucky wigeon
#

Aight

rocky kite
#

Here's a map of good starting resolutions you can use that are compatible with ltx2.3's upscalers and ic lora:

  • 256×384 -> 1.5x -> 384×576 -> 2x -> 768×1152
  • 384×512 -> 1.5x -> 576×768 -> 2x -> 1152×1536
  • 512×768 -> 1.5x -> 768×1152 -> 1.5x -> 1152×1728
  • 512×768 -> 1.5x -> 768×1152 -> 2x -> 1536×2304
  • 512×896 -> 1.5x -> 768×1344 -> 2x -> 1536×2688
  • 512×1024 -> 1.5x -> 768×1536 -> 1.5x -> 1152×2304
  • 512×1024 -> 1.5x -> 768×1536 -> 2x -> 1536×3072
  • 768×1024 -> 1.5x -> 1152×1536 -> 1.5x -> 1728×2304
  • 768×1024 -> 1.5x -> 1152×1536 -> 2x -> 2304×3072
plucky wigeon
#

Trying some videoclip sheninaigans now, I do wanna try the upscaling too

rocky kite
#

note that ltx2.3 ignores whatever upscale value you selected. If you choose the 2x it'll do 2x; if you choose 1.5x it'll do that, so the UI choice doesn't matter

#

I only disable upscaling at all if you set upscale==1

plucky wigeon
#

I was really gonna try downscaling first...

#

Then try upscaling next

rocky kite
#

can't, just gen from small size

plucky wigeon
#

poopee

rocky kite
#

the reference images are down/upscaled automatically for you

plucky wigeon
#

What

rocky kite
#

I do this automatically now

plucky wigeon
#

Then how cant downscale

rocky kite
#

if you start with a 1024x1024 base image, but set video size to 256x256, videostages will automatically downscale the init image to 256x256

#

vs the 3 possible choices in swarmui

plucky wigeon
#

oh, i have to set video resolution 😐

#

I thought this was for that

rocky kite
#

yes, that's what that is, your initial video size

#

I'm just comparing what swarmui native does

plucky wigeon
#

Okay, well this is what I did

rocky kite
#

fuckin' a 👍🏽

plucky wigeon
#

idk did it come out the right size? maybe? lol

rocky kite
#

ok so look at this table #1486751648259772457 message

#

You want your videos to match one of those starter resolutions, either in portrait or landscape mode

#

it tells you your safe paths to upscale to a desired size with ltx2.3

plucky wigeon
#

What bottom thing for

rocky kite
#

Anything other than those resolutions risks running into the upscaler model's problems with "needs to be divisible by 32 at all times", AND the controlnet's requirement of "needs to be divisible by 64 at all times"

plucky wigeon
#

Would be neat if you had these as dropdown options

rocky kite
#

great idea

rocky kite
#

too much as and it bakes it

#

too little and it doesn't apply strongly enough

#

if image is on frame 1, set to strength=1 for stage 0; afterward drop to 0.8/0.6 etc the more stages a clip has

#

less strength gives ltx more freedom to compose a cohesive video, vs being yanked back to that specific image - especially for non-first-frame ref images

plucky wigeon
#

There's something wrong with the parsing'ing

rocky kite
#

what's wrong with it

plucky wigeon
#

base should be sonic, bas2edit should make him spooky (it do), but only video should know he is annihilate

rocky kite
#

gimme the workflow or screenshot your swarm sidebar

plucky wigeon
#

In my brain, this should gen sonic, klein should make it mutate into donal trump, and then the video should be the resultant amagimation apologizing

rocky kite
#

oh I see

#

Needs removed/ignored. Gotta see how swarm does it

plucky wigeon
#

Why steps 20

rocky kite
#

just randomly testing

#

trying to reproduce issue

plucky wigeon
#

Aight

rocky kite
#

btw make sure to install acestepfun to make music videos

plucky wigeon
#

Let me break one thing at a time here.

rocky kite
#

I found issue, fixing

plucky wigeon
#

Is updated and restarted. Trying with correct upscale value (maybe)

rocky kite
#

next up in videostages, support for more of the ltx loras

#

Update both extensions 👍🏽

#

I've learned a lot writing these extensions, videostages is my latest "things I know" project, so all my older extensions need some updating

rocky kite
#

after more lora support is added, then next logical step is doing intra-clip features

plucky wigeon
#

Next up: Do the stages in order.

rocky kite
#

as in, two clips 0 and 1, and clip 1 can reference the last x frames of clip 0 for smooth transitions

rocky kite
plucky wigeon
#

All gens, all edits, all videos. Not swap model 9 times

#

RIP Generation Time: 21.52 sec prep, 7.10 min gen

rocky kite
#

but that's a comfy thing?

#

in my tests if your video depends on refiner stage, everything previous must be executed first

plucky wigeon
#

Maybe, no way we can send to comfy in a diff order?

rocky kite
#

explain

plucky wigeon
#

Dont comfy do it in the order we send it?

rocky kite
#

I'm unsure what algo it uses to figure out the graph dependencies, but it runs in required order, not specified order

#

"B depends on A, so run A first"

#

or, "A depends on Z, so run Z first"

plucky wigeon
#

If we send Images: 6. Can we not send them to Comfy in a diff order or nah

rocky kite
#

oh yeah that's way out of scope of whatever dumb things I do

plucky wigeon
#

Fair, just what popped into my head rn

rocky kite
#

however

#

you can gen the images first, then go back and use them as init images

plucky wigeon
#

Yeah, I can. I just thought I couldnt be the only one who generates a dozen things at once lol

#

esp. if say you trying to do a storyboard in one long mondoprompt

rocky kite
#

that comfy can be smart about

#

in my tests with 4 clips, multiple stages, each with different init images, it genned the images first before moving on to the video portions

plucky wigeon
#

Sonic
<videostage[0]>He jumps
<videostage[1]>He spins
<videostage[2]>He eats chili dog
<videostage[3]>He does another thing

rocky kite
#

videoclip

plucky wigeon
#

Weird, cuz in mine is makes an image, edits it, then video

#

I was making an example

rocky kite
#

yes

plucky wigeon
rocky kite
#

if you want to be nuts about it, I can add a toggle "append to video" to a clip. if unchecked it saves the video clip separately

#

what are you doing baratan

plucky wigeon
#

I am pressing Generate button

rocky kite
#

oh not passing video to z

plucky wigeon
#

yes, no, 6 different gens lol

#

but, loads each model, each time instead of doing the same model 6 times

rocky kite
#

Wonder how you'd implement something like that in swarm

#

each gen has no knowledge of other gens

plucky wigeon
#

You probably right.

rocky kite
#

either that or just ... maybe a "Make da Video" button

plucky wigeon
#

But its one Generate press, so possibly it could parse for model idk

rocky kite
#

so you can gen all your images first, keeping the prompt with <videoclip> intact, but with videostages disabled

rocky kite
#

then once you have your 50 images, enable videostages and go through each image and press da button

plucky wigeon
#

Yeah, and then the videoprompt is used with the image as the init

#

i like it

rocky kite
#

These are things that I don't think about cus ... it doesn't affect me, so it's nice to get feedback on improvements

plucky wigeon
#

Thats why Im the 3kndfort or whatever its called

rocky kite
#

I'll add a button tonight, everything else is already implemented

pallid aurora
pallid aurora
pallid aurora
# rocky kite Wonder how you'd implement something like that in swarm

Generally, not specific to videostages.

I've thought it'd be nice to be able to cache multiple runs of output latents at each stage to reduce the number of times models swap in and out when running, say, five generations.

[Sorta kinda like A1111's old fashioned "how many images in a batch" but instead of in parallel, do it serially to not impact VRAM use by much.]

In the simplest case, assuming all on GPU and maybe GPU VRAM constraints: load clip/whatever, turn five prompts into encodings/latents in serial fashion, cache the 5 new input latents, load primary model, pass each input latent through tensors/networks n steps in serial fashion and cache the five output latents, load VAE, pass each output latent through the VAE and output the five images.

It's possible the current state of "smart memory" VRAM/RAM swapping has mitigated (half/much?) of this for nvivia gpus. Also, I have no idea what the sizes of input/output latents are and if there might be a large amount of VRAM/RAM swapping added for something like this.

delicate crypt
#

If you have enugh RAM then i think Comfy should keep all of the models hot in RAM and copying that over should be fast

#

Tho i had a case where i was doing Klein 9B -> SeedVR2 when i was doing a batch job to clean watermarks from images for training on 200+ images. It did a lot of model reloading for each image. In theory it could do it without reloading by running Klein 9B on GPU0 and then doing SeedVR2 on GPU1 GPU2 GPU3 (since that step is slower). But i could imagine getting that to work inside Swarm would likely need all sorts of crazy internals reworking

#

But if i was less dumb i could first do a batch for Klein 9B into a temp folder, then feed that whole flder trough SeedVR2 in a second batch job

pallid aurora
#

Yeah, in that case I might do Batch process on dir1 with klein, output to dir2. then batch process on dir2 with seedvr2.

#

heh, yeah.

delicate crypt
#

But since i got 4 GPUs it is a speed cheatcode so i don't realy worry about batches as much It does i think 4 sec per image average on Klein 9B

#

The thing i want is working previews for LTX2. Come on Comfy team get off your ass

rocky kite
rocky kite
#

I'm working on the dimensions dropdown ... here's the final, VideoStages-approved list:

* 192×256 ⇑ 1.5x             ⇒ 288×384
* 192×256 ⇑ 1.5x ⇑ 1.5x      ⇒ 432×576
* 192×256 ⇑ 1.5x ⇑ 2x        ⇒ 576×768
* 192×256 ⇑ 1.5x ⇑ 1.5x ⇑ 2x ⇒ 864×1152

* 256×384 ⇑ 1.5x             ⇒ 384×576
* 256×384 ⇑ 1.5x ⇑ 1.5x      ⇒ 576×864
* 256×384 ⇑ 1.5x ⇑ 2x        ⇒ 768×1152   # safe for controlnet
* 256×384 ⇑ 1.5x ⇑ 1.5x ⇑ 2x ⇒ 1152×1728

* 384×512 ⇑ 1.5x             ⇒ 576×768
* 384×512 ⇑ 1.5x ⇑ 1.5x      ⇒ 864×1152
* 384×512 ⇑ 1.5x ⇑ 2x        ⇒ 1152×1536  # safe for controlnet
* 384×512 ⇑ 1.5x ⇑ 1.5x ⇑ 2x ⇒ 1728×2304

* 384×640 ⇑ 1.5x             ⇒ 576×960
* 384×640 ⇑ 1.5x ⇑ 1.5x      ⇒ 864×1440
* 384×640 ⇑ 1.5x ⇑ 2x        ⇒ 1152×1920
* 384×640 ⇑ 1.5x ⇑ 1.5x ⇑ 2x ⇒ 1728×2880

* 512×768 ⇑ 1.5x             ⇒ 768×1152
* 512×768 ⇑ 1.5x ⇑ 1.5x      ⇒ 1152×1728  # safe for controlnet
* 512×768 ⇑ 1.5x ⇑ 2x        ⇒ 1536×2304  # safe for controlnet
* 512×768 ⇑ 1.5x ⇑ 1.5x ⇑ 2x ⇒ 2304×3456

* 512×896 ⇑ 1.5x ⇑ 2x        ⇒ 1536×2688  # safe for controlnet

* 512×1024 ⇑ 1.5x ⇑ 1.5x     ⇒ 1152×2304  # safe for controlnet
* 512×1024 ⇑ 1.5x ⇑ 2x       ⇒ 1536×3072  # safe for controlnet

* 768×1024 ⇑ 1.5x ⇑ 1.5x     ⇒ 1728×2304  # safe for controlnet
* 768×1024 ⇑ 1.5x ⇑ 2x       ⇒ 2304×3072  # safe for controlnet
#

and obviously can reverse the width+height

rocky kite
#

dimensions dropdown added

plucky wigeon
#

Hmm... my gens keep seeing the videoclip in the previous stages

rocky kite
#

Did you update base2edit

plucky wigeon
#

I updated everything when I got home today.

rocky kite
#

gimme da prompt

plucky wigeon
#

It was

Galadriel /(The Lord of the Rings/)
<edit>reskin this into a 35mm film (Kodak Portra 400) photo.
<videoclip>She speaks and says "Fought in the Clone Wars? Luke, I started the Clone Wars and I fucking finished them."```
#

oh nvm its not updated for some reason. 🙁

rocky kite
#

oh

plucky wigeon
#

Its not updated. I lied I guess.

plucky wigeon
plucky wigeon
rocky kite
#

whatcha mean

plucky wigeon
#

Wether I generate the init images first or not, how do I make a video using two different ones?

Galadriel /(The Lord of the Rings/)
<edit>reskin this into a 35mm film (Kodak Portra 400) photo.
<videoclip>She speaks and says "One does not simply walk into Mordor, Frodo."

Donal Trump
<edit>The photo was processed to a realistic style, creating a lifelike photo. Transform setting into realism style. balanced contrast. remove extra motion lines. remove speech bubbles. Remove text. bright professional studio lighting.
<videoclip>He speaks and says "Fought in the Clone Wars? Luke, I started the Clone Wars and I fucking finished them."```
#

I guess I would manually upload the inits in the videoclip options and then use <videoclip[0] etc?

rocky kite
#

yeah that works

plucky wigeon
#

Hmm, even though I provided images it still generated one with z-image anyway?

rocky kite
#

yeah that's disconnected

plucky wigeon
#

Do you intend for the metadata to contain all of this?

rocky kite
#

oh shit haha

plucky wigeon
#

Well, poop.

rocky kite
plucky wigeon
rocky kite
#

switch to text-to-video mode should help

#

or keep as-is and set image resolution to like 256

plucky wigeon
#

lol it makes rando golf lady every time

rocky kite
#

oh but it would still load haha

plucky wigeon
#

How well does this work to extend?

river slate
#

finally getting my hands on VideoStage and trying to figure out which option does what

#

so far 1 finding: in a 2 stage workflow with LTX 2.3 (single clip), there's no LTXVPreprocess node (the weird one that degrade a bit the image quality so LTX makes more motion) for the refiner stage, is it intentional?

#

and is there a reason for the refiner reference image strength to be at a 0.8 default?

#

and is it possible to have the last frame of clip 0 to be the first frame of clip 1? I can't find a way to chain them 🤔

plucky wigeon
#

I like Hippotes suggestions/feedback.

rocky kite
rocky kite
#

if you don't have an upscale node it reuses the already-existing image

river slate
rocky kite
#

Maybe making it a prompt-only <param> thing? The UI is getting more and more crowded

river slate
#

in the UI it would only expand the ref image dropdown

rocky kite
#

frame is already there though

#

but that's for clip-level (ie, all stages)

#

you want a stage-level frame slider

river slate
#

no

#

say I want 2 x 20s chained clips but I only have the first image, I would like the last frame of clip0 to be the first frame of clip1

rocky kite
#

ah right, referencing the actual video, not the ref image

#

yeah that's still not in yet. I thought the non-native transitions lora would help with that, but apparently not

river slate
#

there's transition lora? 😮

rocky kite
#

because it's not just "last frame of video clip n", it's "last x frames of video clip n"

#

but this also requires referencing pre-upscale video pixels, not just "the last clip"

#

consider:

  • clip 0: stage 0, 512x512 -> upscale 2x -> stage 1, 1024x1024
  • clip 1: stage 0, 512x512 using last 5 frames from clip 0
#

I guess we could just downscale like we're already doing

#

I automatically resize/crop ref images to the correct expected dimensions:

#

and I've already implemented batch rescaling for controlnet support; this would just be another identical idea

#

👍🏽

#

I'll add it to my todo for the week; after I see if ComfyTyped project is viable (I think it is, I've already rewritten to use it, and I'm kind of excited by it)

plucky wigeon
#

Usage:

On the GitHub would be nice. I know you explained everything and even posted pictures in here, but it's gonna be impossible for anyone not in the know to DL and use this. From a rando user perspective anyway.

eager crown
#

NGL...I'm lost on video clip, video stage and how to change the prompts in between. Also, has anyone used this with wan 2.1 or 2.2 with good results? Thank you.

rocky kite
#

yes I need docs

plucky wigeon
#

I'd like to suggest just start with the basic # installation, # usage format that similar projects use and give a few examples on how to use it. Helps how it works click in peoples heads.

#
# Usage
## Basic Video Generation
### Text-to-Video
### Image-to-Video
### Video-to-Video?

## Upscaled Video Generation
### Text-to-Video
### Image-to-Video
### Video-to-Video?

## Extended Video Generation
### Text-to-Video
### Image-to-Video
### Video-to-Video?

## Advanced Usage
### Using supplied external audio
### What is AceFunStepDance
rocky kite
#

I'm happy to take suggestions. I'm about to fire up the bbq grill and cook up some tbones

river slate
#

the Length (seconds) slider do nothing? with a simple 1 clip 1 stage + ref image from upload I have to set the frame count from Text2Video section to have them change in the workflow
edit: this is weird, I switch to another init image and now the Length takes over the frame count

#

and I feel like I'm missing the point of the stage skip feature 🤔
If I have a 3 stage workflow, can I see the intermediate outputs (so I can cancel it after first stage if I don't like it)?

rocky kite
#

So what I do is I set up my 3 stages, disable the first two then gen 50 tiny videos videos

#

I delete the ones I don't like then re-enable the other two stages and upscale the ones I kept

#

now I don't have to recreate the second two stages, just two button clicks

#

or if you set your seet to a specific value, you can gen first stage; if you like it, activate second stage and regen, only second gen kicks in, the rest is cached

river slate
rocky kite
plucky wigeon
# rocky kite

Can you add square ones, or is squares bad for LTX-2.3?

rocky kite
#

No, I guess all of those squared would be just fine

river slate
#

the workflow generated with multiple reference images is broken :x

#

it ends like that with a CropGuides node (idk what it does) and no audio decode node (<- that's what really fail to run)

rocky kite
#

cropguides without a downstream link is correct

#

the no audio thing is of interest to me

river slate
#

basic single stage, native audio, only 4 reference images at regular frames spacing, all other things default

rocky kite
#

can you take a screenshot of the videostages section? I'm in the middle of a large refactor pending a merge into swarmui core, so it might be a day before I can fix it, but I sure can't reproduce in any of my regular workflows

river slate
#

and there's still something weird with frame number: sometimes it gets them from Text2Video > Frames, sometimes from Clip > Length, I can't figure out why so I just set both but weird

rocky kite
#

it's probably that eros model? We're hooking into swarm for decision if a model supports audio

#

look earlier in the graph, is the audio being applied in earlier stages?

river slate
#

yeah first half of it look right with Audio VAE loader > LTXV Empty Latent Audio > LTXVConcatAVLatent > Ksampler

#

oh wait with only 2 ref images it's fine

#

break at 3 🤔

#

and once it's borked, going back to 2 doesn't fix it

#

only 1 ref image as first frame is fine tho

#

OH GOT IT

#

it borks as soon as Frame is not 1

#

this runs, we see the audio decode node at the end (but stupid as both frame are at 1)

#

this doesn't

river slate
#
<videoclip[0,1]><lora:LTX-2.3/ltx-2.3-22b-distilled-lora-1.1_fro90_ceil72_condsafe:0.5>```
this is very nice syntax 🥰
rocky kite
#

"what could this look like so I don't hate it?"

river slate
#

OH

#

it doesn't carry global prompt tho

#

I was shoving it at the end of the prompt but i'll have to put the prompt in a variable

rocky kite
#

👀 <b2eprompt[global]>

#

but that won't kick in if you don't have b2edit enabled, so, sucks

river slate
#

leaving it here but it's between VideoStages and SeedVR: upscaling with SeedVR a video generated from 2 stage with upscaling reuse the base resolution instead of the 2nd stage resolution

#

I have Stage 0 (512) -- x2 --> Stage 1 (1024) --> SeedVR x1.5 --> hoped for 1536, got 768

rocky kite
#

yeah I had to change the resolution in seedvr

river slate
#

oh right I can set it manually instead of the multiplier

pallid aurora
# river slate ```<videoclip[0,0]><lora:LTX-2.3/ltx-2.3-22b-distilled-lora-1.1_fro90_ceil72_con...

I kind of wish that SwarmUI didn't send the newline right before a segment marker, as some text-input models (clips,llms,magic) will give different output inputs that include the same text but differeng pre/post text whitespace.

I've started ending segments like this:
<base>
blah blah blah <comment:

<refiner/video/whatev>

so that I don't get lost scrolling up and down. :/

Maybe a whitespace policy setting or control would be nice, so as to not break current param set reproducibility (as much as that can be done with a changing code base or five).

rocky kite
#

Coming soon:

  • use controlnet videos in ltx and wan
  • use controlnet audio in your videos
pallid aurora
#

Quick Q Juan - if I wanted to do one-shot Text-to-Image to SeedVR2-to-Image into the front of the pipeline of VideoStages...is there a way to tell VideoStages to grab it from there? When I use SeedVR2 with "Before Video" selected (to force image-based upscale), it works! But only after the entire VideoStages pipeline ends, lol.

#

I left the option in Videostages set to (output of ) Refiner but I'm not sure currently this drop down gives one that works for the above?

rocky kite
pallid aurora
#

Ah ok. Thanks for the clarification.

pallid aurora
#

I wonder if SwarmUI needs an order-of-addons(?) or order-of-nodes introspection to be able to adjust things like "where to grab input from" or similar? If it doesn't already.

plucky wigeon
#

THe grid generator has the same problem with upscales and edits too

pallid aurora
#

Like "I am registering that I take in the following configurations: image, image + ref_image(max of n), ref_image (max of n), video, video+audio, video + ref_image, or video + audio + ref_image". I can output: image, video, or video + audio.

#

Oh and assume the input list has with and without text variants.

#

I dunno, maybe getting silly.

rocky kite
#

big update pushed

#

you need to update swarmui first

river slate
#

sooo updating went kinda fine, like no regression from previous version

river slate
rocky kite
rocky kite
river slate
#

try to generate or import workflow in Comfy tab, error

rocky kite
#

I've got a million tests and have run through that scenario a million times

#

this works fine. what else is up with your settings then

river slate
#

I have image source upload instead of from refiner

rocky kite
#

so text-to-video?

river slate
#

what? no

#

i2v

rocky kite
#

can you paste all your settings? This is still working for me

river slate
#

if I remove the 2nd reference or set it as Frame 1 it generates a valid workflow

rocky kite
#

yeah that's weird as hell. I can't actually hit it

#

Can you send payload of the import from generate tab button?

river slate
#

(heavy cause images)

rocky kite
#

the damn download button doesn't work, damn you discord

#

can you zip it up?

river slate
#

(I can't share the payload from "import" button as it's truncated to 1M in dev tools)

#

(or brb finding smol images)

rocky kite
#

nah you're fine

river slate
rocky kite
#

oh are you using a different ltx checkpoint?

river slate
#

this was with official distilled-1.1 bf16

#

but I have exact same thing with official dev

#

and my nvfp4 ones

rocky kite
#

ok try updating

river slate
#

it works \o/

#

at least import, generate time

rocky kite
#

gonna try to get the frame picker out today, but that's in #1462288362537746574

river slate
#

especially creepy with empty prompt but the ref frames are where they are supposed to

rocky kite
#

wooo

#

ugh, acestepfun not connecting. How'd that get missed

river slate
#

claude bad booooh 😋

rocky kite
#

nah that was probably me

#

the rewrite was mostly me with claude helping out for syntax

#

although it's got me thinking, would be pretty neat to connect my FOSS projects github issues to a running claude instance that auto-fixes bugs as they are reported

rocky kite
#

that booty got its own soundtrack

pallid aurora
#

the only trend i've noticed is you either get good movement or good sound, usually not both, womp womp.

pallid aurora
#

@rocky kite - have you had success with VideoStages with Sulphur2Base? If so, pretty standard? I'm about to go in, lol.

Also, I'm wondering if the fact that the Swarm Internal "Output Intermediate Images" doesn't actually capture the image before SeedVR2 means that maybe the point at which SeedVR2 injects itself isn't being done correctly? Thoughts?

rocky kite
#

Let me know how it does

pallid aurora
#

when i gen small area T2V videos with Sulphur Base...oh boy it's rather trash...even using the suggested sampler/scheduler. So I guess I'll try, but I have a bad feeling about it. Looks like the official workflow has much different sampler/scheduler settings for each layer, but starts with large area, so maybe it's an ill fit for juan-small-to-big-workflow.

#

But yeah, I'll tinker.

rocky kite
#

T2I like image to video?

#

Or straight from text

pallid aurora
#

sorry, T2V

#

edit, fixed above.

rocky kite
#

Ltx doesn’t do well with t2v from my attempts

#

Like try generating a single image like with wan

pallid aurora
#

Yeah, apparently Sulphur was trained specifically for T2V and 10eros is a modified merge to integrate support for I2V. So, for I2V use 10Eros. For T2V use Sulphur. According to the trainers/mergers.

#

So, like, you can generate 201+ frame meh videos at large area using T2V with sulphur Base. Just need to experiment to see if it's actually useful without using the complex confyui workflow they shipped with it. If you are very patient.

#

You can get better looking videos using the distill, but that loses most of the flexibility of sulphur training.

#

Clarified above that I'm talking about Sulphur Base.

rocky kite
#

Enjoy!

rocky kite
# rocky kite Enjoy!

You can generate a video with n stages. Then you can select it in history and press the "Refine Video" button. It works just like the Refine Image button where the current prompt and all other settings are used, only the original seed is kept - AND the selected video is used and NOT regenerated.

You must define at least n+ 1 stages for clip 1 for this to work.

In other words, if you gen a video with only stage0, you must define at least stage0 + stage1; stage0 is SKIPPED. If you gen with stage0 and stage1, you must define at least stage0 + stage1 + stage2.

rocky kite
#

If you don't want to regenerate images, add them as uploaded ref images

river slate
#

something broke

#

SwarmUI.Builtin_ComfyUIBackend.WGNodeData.FPS
Swarm & VideoStage are up to date, Comfy is not (from before the memory fuckery)

rocky kite
#

need to rebuild swarm

#

should have happened automatically on update

river slate
#

I believe it did, and I just restarted from the launch-dev script and same error :x

#

(I only removed the b64 image data for the filesize)

rocky kite
#

I don't, it's a c# change, not workflow

#

Can you uninstall and reinstall the extension? Something's not triggering the latest commit

river slate
#

I just did

Already up to date.```
but sure I can reinstall
#

19:40:02.485 [Error] [WebAPI] Error handling API request '/API/UninstallExtension' for user 'local': Extension deletion failed, you will need to manually delete the extension folder from inside SwarmUI/src/Extensions

that was unexpected 🤣

#

'k manually delete did it, restarting..

#

nope, still same error

rocky kite
#

oh shit comfytyped needs updated

#

yeah give me one minute

rocky kite
river slate
#

hmm yeah sooo... workflow generates and run without error, but reference image (via image upload) is ignored :x

#

it does t2v only 😅

#

wait no

#

idk what happened my ref was ignored but F5 fixed it

#

yep, everything fine again, ty 😄

rocky kite
#

For my future reference, what spectrum settings worked for me (with a patch), for ltx 30 steps, 4 cfg

odd musk
rocky kite
rocky kite
#

Pushed support to videostages for lanczos and other upscalers

#

and I'm working on adding support for seedvr2

river slate
#

how different from existing seedvr2?

rocky kite
#

basically like any other upscaler that videostages already uses; seedvr2 has no concept of videostages' multiple clips/stages, and it's "before video" and "after video" end up at the same location: end of workflow

river slate
#

oooh seedvr between stages

rocky kite
#

So, my idea: let seedvr2 add its node at the very end, then videostages comes in and moves it to the right location(s)

river slate
#

that gonna be painfully slow but why not

rocky kite
#

seedvr is ... faster than ltx's upscaler for me

river slate
#

what how?

rocky kite
#

I thought this was common for everyone?

river slate
#

a 640p -> 960p gen is like 3 min for me with latent upscaler x1.5, then putting it through seedvr without upscale (aka downscale 0.5 and output 960p again) is 6 min alone

#

with 3B fp16, but vae encode/decode is more than half the time anyway

rocky kite
#

I do utilize my gpu to its maximum though

#

I'm hopeful that this means the videostage swarmksampler refiner after seedvr upscaling means far less steps are needed, too

#

maybe as low as 1 or 2 steps just to help lessen seedvr's plasticisity

wary epoch
#

@Juan Can one by pass stage 0 and use an init image instead of genning one?

river slate
wary epoch
#

Ya trying to get to grips with it in a first frame last frame workflow. Will the transition lora work here as well?

rocky kite
#

it should

#

if it's just a regular lora

wary epoch
#

Ya its the one you mentioned above in the chat.

rocky kite
#

right. I've been juggling so many balls recently I forget what I've done

#

for last frame just check the box for counting backwards and set it to 1

#

so no matter the length of your video it'll always be -1

wary epoch
#

Cool. How do I stop an image from being generated as base and to just use the supplied ref image? Just turn control to 0 ?

rocky kite
#

yup

#

You can then see in the comfy workflow that it's disconnected

wary epoch
#

Can an existing video be dropped into the workflow for editing?

rocky kite
#

click "refine video" and it'll jump to the second stage

wary epoch
#

In terms of refine vid what does it do?

rocky kite
#

it jumps to the second stage

#

it takes the video as-is and runs it through clip0.stage1 (second stage)

#

and whatever other stages you have defined after

#

this is useful if you gen a low-rez video with stage1+ disabled; you can then enable them and click refine video and it won't regen the base video again, it'll just jump straight to stage1

#

well it'll still regen other things it needs like images or audio

#

because those are baked into the video, right

wary epoch
#

@rocky kite Fuck me Juan this is cool. Kudos dude. Quick question re dimensions if I use a 9:16 init image and I'm using the lowest dimension and up scaling x 2, how do I maintain the 9:16 aspect ratio?

#

Seems to be cutting it off

rocky kite
#

set the video dimensions

#

but note that if u sing ltx2.3, it really wants dimensions divisible by 32

#

if you're going to use a dimension not in the list in the dropdown, you'll want to switch to lanczos upscaling, otherwise the ltx upscaler will add shit to the borders

#

or seedvr2

wary epoch
#

Oh ok will give it a go.

#

Fuck me I dunno why I wasn't upscaling sooner reduces gen times considerably

rocky kite
#

I've been blabbing about this for months!

wary epoch
#

Heh I was following your extension development closely, but I was too overwhelmed to give it a proper go. Getting the hang of it now...

#

288x512 is my sweet spot 30 sec for a 7 sec vid with upscale. Oh and can we use multiple loras?

rocky kite
#

yes, it's additive

#
<video>applies to all videos
<videostage>applies to all videostages videos
<videostage[0]>applies to all videostage clip0
<videostage[0,0]>applies to videostage clip0.stage0
<videostage[0,1]>applies to videostage clip0.stage1
wary epoch
#

the above are global prompts I take it? How do I apply other loras do I do them in brackets ya? Thw way we first did with Wan?

rocky kite
wary epoch
#

Hmm How do I stop it from genning a base img with Z even though its not used slows inference

wary epoch
#

The one I have been initially using for some reason it now gens a base img even though its not used

rocky kite
#

I'm not following though

wary epoch
#

So I was or so I though bypassing stage 0 as I was using my own uploaded images, but when i run it it first gens a z image base image which is wholly random then it happily fulfills everything else slows it down a bit is all

#

Wholly random img it generates weird

rocky kite
#

Well do you have zimage selected as the main model

wary epoch
#

yup

rocky kite
#

right so it's doing the text-to-image, image-to-video workflow

#

select ltx as the main model

wary epoch
#

cool