LTX-2.5 22B dev generating synchronized audio and video in one sampling pass per stage, on the shape of the ComfyUI LTX-2.5 templates: text or image to video in two stages (half resolution, 2x latent upscale, a short distilled refine), first/last frame and reference-sheet (Ingredients IC-LoRA) in one stage, and an optional 2x pixel re-render through the Pixel Spatial Upscaler IC-LoRA.
Tags: motionltx-2.5audio-videotext-to-videoimage-to-videofirst-last-framereference-sheetrunpod-serverlessvolume-backed
Inputs (23)#
The typed parameter surface callers bind when they request this workflow. Enum options and numeric bounds are the values the workflow document declares.
modeenumdefault t2vt2vi2vflf2vingredientspromptstringrequireddefault a wooden boat creaking against a dock at dawn, gulls calling in the distancenegative_promptstringdefault pc game, console game, video game, cartoon, childish, uglyshotsarraydefault []shot_transitionenumdefault hard_cuthard_cutmatch_cutdissolvestart_imageimageend_imageimageimage_strengthfloatdefault 0.7min 0.0max 1.0reference_sheetimageingredients_strengthfloatdefault 1.0min 0.0max 2.0generate_audiobooleandefault trueresolutionenumdefault 1280x7041280x704704x12801024x576768x4481920x1088lengthintegerdefault 121min 9max 193auto_durationbooleandefault falsefpsenumdefault 242425stepsintegerdefault 30min 1max 40cfg_videofloatdefault 3.0min 0.0max 20.0cfg_audiofloatdefault 7.0min 0.0max 20.0seedintegerdefault 0min 0samplerenumdefault eulereulerdpmpp_2muni_pceuler_ancestralschedulerenumdefault simplesimplenormalbetakarrasspeed_modeenumdefault defaultdefaultdistilledupscaleenumdefault nonenonepixel_x2ComfyUI node graph (63)#
The executable ComfyUI prompt graph: 63 nodes across 35 distinct node classes, wired by 85 data dependencies. Nodes tinted green come from a custom node pack this workflow declares; the rest are ComfyUI core / baked-community classes.
Nodes (63)#
1UNETLoadercoreunet_name = ltx-2.5-22b-dev-transformer-bf16.safetensorsweight_dtype = defaultMODEL2CLIPLoadercoreclip_name = gemma4-12b-with-proj-ltx-2.5-bf16.safetensorstype = ltxvdevice = defaultCLIP3VAELoadercorevae_name = ltx-2.5-audio-vae-bf16.safetensorsVAE4VAELoadercorevae_name = ltx-2.5-video-vae-bf16.safetensorsVAE5LatentUpscaleModelLoadercoremodel_name = ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensorsLATENT_UPSCALE_MODEL6LoraLoaderModelOnlycoremodel = ◂ node 1 · out[0]lora_name = ltx-2.5-22b-distilled-lora-450-bf16.safetensorsstrength_model = 1.0MODEL7CLIPTextEncodecoreclip = ◂ node 2 · out[0]text = {{prompt}} tmplCONDITIONING8CLIPTextEncodecoreclip = ◂ node 2 · out[0]text = {{negative_prompt}} tmplCONDITIONING9LTXVConditioningcorepositive = ◂ node 7 · out[0]negative = ◂ node 8 · out[0]frame_rate = {{float(fps)}} tmplCONDITIONINGCONDITIONING10EmptyLTXVLatentVideocorewidth = {{ingredients_size.width if mode == 'ingredients' else resolution_map[resolution].width // stage_one_divisor[mode]}} tmplheight = {{ingredients_size.height if mode == 'ingredients' else resolution_map[resolution].height // stage_one_divisor[mode]}} tmpllength = {{length}} tmplbatch_size = 1LATENT11LoadImagecoreimage = {{start_image}} tmplIMAGEMASK12ResizeImageMaskNodecoreinput = ◂ node 11 · out[0]resize_type = scale longer dimensionresize_type.longer_size = 1536scale_method = lanczosIMAGE13LTXVPreprocesscoreimage = ◂ node 12 · out[0]img_compression = 18IMAGE14LTXVImgToVideoInplacecorevae = ◂ node 4 · out[0]image = ◂ node 13 · out[0]latent = ◂ node 10 · out[0]strength = {{image_strength}} tmplbypass = falseLATENT15LoadImagecoreimage = {{end_image}} tmplIMAGEMASK16ResizeImageMaskNodecoreinput = ◂ node 11 · out[0]resize_type = scale dimensionsresize_type.width = {{resolution_map[resolution].width}} tmplresize_type.height = {{resolution_map[resolution].height}} tmplresize_type.crop = centerscale_method = nearest-exactIMAGE17ResizeImageMaskNodecoreinput = ◂ node 15 · out[0]resize_type = scale dimensionsresize_type.width = {{resolution_map[resolution].width}} tmplresize_type.height = {{resolution_map[resolution].height}} tmplresize_type.crop = centerscale_method = nearest-exactIMAGE18LTXVPreprocesscoreimage = ◂ node 16 · out[0]img_compression = 18IMAGE19LTXVPreprocesscoreimage = ◂ node 17 · out[0]img_compression = 18IMAGE20LTXVAddGuidecorepositive = ◂ node 9 · out[0]negative = ◂ node 9 · out[1]vae = ◂ node 4 · out[0]latent = ◂ node 10 · out[0]image = ◂ node 18 · out[0]frame_idx = 0strength = {{image_strength}} tmplCONDITIONINGCONDITIONINGLATENT21LTXVAddGuidecorepositive = ◂ node 20 · out[0]negative = ◂ node 20 · out[1]vae = ◂ node 4 · out[0]latent = ◂ node 20 · out[2]image = ◂ node 19 · out[0]frame_idx = -1strength = {{image_strength}} tmplCONDITIONINGCONDITIONINGLATENT22LoadImagecoreimage = {{reference_sheet}} tmplIMAGEMASK23ResizeAndPadImagecoreimage = ◂ node 22 · out[0]target_width = {{ingredients_size.width}} tmpltarget_height = {{ingredients_size.height}} tmplpadding_color = blackinterpolation = lanczosIMAGE24RepeatImageBatchcoreimage = ◂ node 23 · out[0]amount = {{max(length, 121)}} tmplIMAGE25LTXICLoRALoaderModelOnlycoremodel = {{model_source_map[speed_mode]}} tmpllora_name = ltx-2.5-22b-ic-lora-ingredients-0.9.safetensorsstrength_model = {{ingredients_strength}} tmplMODELFLOAT26LTXAddVideoICLoRAGuidecorepositive = ◂ node 9 · out[0]negative = ◂ node 9 · out[1]vae = ◂ node 4 · out[0]latent = ◂ node 10 · out[0]image = ◂ node 24 · out[0]frame_idx = 0strength = 1.0latent_downscale_factor = ◂ node 25 · out[1]crop = disableduse_tiled_encode = falsetile_size = 256tile_overlap = 64CONDITIONINGCONDITIONINGLATENT27LTXVEmptyLatentAudiocoreaudio_vae = ◂ node 3 · out[0]frames_number = {{length}} tmplframe_rate = {{float(fps)}} tmplbatch_size = 1LATENT28LTXVConcatAVLatentcorevideo_latent = {{video_latent_source_map[mode]}} tmplaudio_latent = ◂ node 27 · out[0]LATENT29LTXVDualCFGGuidercoremodel = {{model_source_map['ingredients' if mode == 'ingredients' else speed_mode]}} tmplpositive = {{positive_source_map[mode]}} tmplnegative = {{negative_source_map[mode]}} tmplvideo_cfg = {{1.0 if speed_mode == 'distilled' else cfg_video}} tmplaudio_cfg = {{1.0 if speed_mode == 'distilled' else cfg_audio}} tmplGUIDER30KSamplerSelectcoresampler_name = {{distilled_sampler_map[mode] if speed_mode == 'distilled' else sampler}} tmplSAMPLER31SamplerEulerAncestralcoreeta = 0.0s_noise = 1.0SAMPLER32BasicSchedulercoremodel = ◂ node 1 · out[0]scheduler = {{scheduler}} tmplsteps = {{steps}} tmpldenoise = 1.0SIGMAS33ManualSigmascoresigmas = 1.0, 0.99375, 0.9875, 0.98125, 0.975, 0.909375, 0.725, 0.421875, 0.0SIGMAS34RandomNoisecorenoise_seed = {{seed}} tmplNOISE35SamplerCustomAdvancedcorenoise = ◂ node 34 · out[0]guider = ◂ node 29 · out[0]sampler = {{sampler_source_map['eta0' if mode == 'flf2v' and speed_mode == 'distilled' else 'select']}} tmplsigmas = {{sigma_source_map[speed_mode]}} tmpllatent_image = ◂ node 28 · out[0]LATENTLATENT36LTXVSeparateAVLatentcoreav_latent = ◂ node 35 · out[0]LATENTLATENT37LTXVCropGuidescorepositive = {{positive_source_map[mode]}} tmplnegative = {{negative_source_map[mode]}} tmpllatent = ◂ node 36 · out[0]CONDITIONINGCONDITIONINGLATENT38LTXVLatentUpsamplercoresamples = ◂ node 36 · out[0]upscale_model = ◂ node 5 · out[0]vae = ◂ node 4 · out[0]LATENT39LTXVImgToVideoInplacecorevae = ◂ node 4 · out[0]image = ◂ node 13 · out[0]latent = ◂ node 38 · out[0]strength = 1.0bypass = falseLATENT40LTXVConcatAVLatentcorevideo_latent = {{stage_two_video_source_map[mode]}} tmplaudio_latent = ◂ node 36 · out[1]LATENT41RandomNoisecorenoise_seed = 42NOISE42LTXVDualCFGGuidercoremodel = ◂ node 6 · out[0]positive = ◂ node 9 · out[0]negative = ◂ node 9 · out[1]video_cfg = 1.0audio_cfg = 1.0GUIDER43KSamplerSelectcoresampler_name = euler_ancestralSAMPLER44ManualSigmascoresigmas = 0.85, 0.7250, 0.4219, 0.0SIGMAS45SamplerCustomAdvancedcorenoise = ◂ node 41 · out[0]guider = ◂ node 42 · out[0]sampler = ◂ node 43 · out[0]sigmas = ◂ node 44 · out[0]latent_image = ◂ node 40 · out[0]LATENTLATENT46LTXVSeparateAVLatentcoreav_latent = ◂ node 45 · out[0]LATENTLATENT47VAEDecodeTiledcoresamples = {{video_decode_source_map[mode]}} tmplvae = ◂ node 4 · out[0]tile_size = 512overlap = 64temporal_size = 64temporal_overlap = 16IMAGE48LTXVAudioVAEDecodecoresamples = {{audio_latent_source_map[mode]}} tmplaudio_vae = ◂ node 3 · out[0]AUDIO49LTXICLoRALoaderModelOnlycoremodel = ◂ node 6 · out[0]lora_name = ltx-2.5-22b-ic-lora-pixel-spatial-upscaler-x2-1.0.safetensorsstrength_model = 1.0MODELFLOAT50EmptyLTXVLatentVideocorewidth = {{2 * (ingredients_size.width if mode == 'ingredients' else resolution_map[resolution].width)}} tmplheight = {{2 * (ingredients_size.height if mode == 'ingredients' else resolution_map[resolution].height)}} tmpllength = {{length}} tmplbatch_size = 1LATENT51LTXAddVideoICLoRAGuidecorepositive = ◂ node 9 · out[0]negative = ◂ node 9 · out[1]vae = ◂ node 4 · out[0]latent = ◂ node 50 · out[0]image = ◂ node 47 · out[0]frame_idx = 0strength = 1.0latent_downscale_factor = ◂ node 49 · out[1]crop = disableduse_tiled_encode = falsetile_size = 256tile_overlap = 64CONDITIONINGCONDITIONINGLATENT52LTXVSetAudioRefTokenscorepositive = ◂ node 9 · out[0]negative = ◂ node 9 · out[1]audio_latent = {{audio_latent_source_map[mode]}} tmplCONDITIONINGCONDITIONINGLATENT53LTXVConcatAVLatentcorevideo_latent = ◂ node 51 · out[2]audio_latent = ◂ node 52 · out[2]LATENT54RandomNoisecorenoise_seed = 42NOISE55CFGGuidercoremodel = ◂ node 49 · out[0]positive = ◂ node 51 · out[0]negative = ◂ node 51 · out[1]cfg = 1.0GUIDER56KSamplerSelectcoresampler_name = euler_ancestralSAMPLER57ManualSigmascoresigmas = 1.0, 0.99375, 0.9875, 0.98125, 0.975, 0.909375, 0.725, 0.421875, 0.0SIGMAS58SamplerCustomAdvancedcorenoise = ◂ node 54 · out[0]guider = ◂ node 55 · out[0]sampler = ◂ node 56 · out[0]sigmas = ◂ node 57 · out[0]latent_image = ◂ node 53 · out[0]LATENTLATENT59LTXVSeparateAVLatentcoreav_latent = ◂ node 58 · out[0]LATENTLATENT60LTXVCropGuidescorepositive = ◂ node 51 · out[0]negative = ◂ node 51 · out[1]latent = ◂ node 59 · out[0]CONDITIONINGCONDITIONINGLATENT61VAEDecodeTiledcoresamples = ◂ node 60 · out[2]vae = ◂ node 4 · out[0]tile_size = 512overlap = 64temporal_size = 64temporal_overlap = 16IMAGE62CreateVideocoreimages = {{final_images_map[upscale]}} tmplfps = {{float(fps)}} tmplaudio = ◂ node 48 · out[0]VIDEO63SaveVideocorevideo = ◂ node 62 · out[0]filename_prefix = ltx25-avformat = mp4Parameter banks (16)#
The prompt / configuration lookup tables this workflow keys into from its inputs — the vocabulary that turns a style / palette / preset selection into graph parameters.
resolution_map (5)#
1280x704{"width": 1280, "height": 704}704x1280{"width": 704, "height": 1280}1024x576{"width": 1024, "height": 576}768x448{"width": 768, "height": 448}1920x1088{"width": 1920, "height": 1088}ingredients_size (2)#
width768height448stage_one_divisor (4)#
t2v2i2v2flf2v1ingredients1requires_families (1)#
ltx25model_source_map (3)#
default10distilled60ingredients250positive_source_map (4)#
t2v90i2v90flf2v210ingredients260negative_source_map (4)#
t2v91i2v91flf2v211ingredients261video_latent_source_map (4)#
t2v100i2v140flf2v212ingredients262stage_two_video_source_map (2)#
t2v380i2v390distilled_sampler_map (3)#
t2veuler_ancestrali2veuler_ancestralingredientseuler_ancestral_cfg_ppsampler_source_map (2)#
select300eta0310sigma_source_map (2)#
default320distilled330video_decode_source_map (4)#
t2v460i2v460flf2v372ingredients372audio_latent_source_map (4)#
t2v461i2v461flf2v361ingredients361final_images_map (2)#
none470pixel_x2610post_render (1)#
ltx25_shots_duration{"shots_input": "shots", "transition_input": "shot_transition", "prompt_input": "prompt", "mode_input": "mode", "auto_duration_input": "auto_duration", "length_input": "length", "fps_input": "fps", "duration_modes": ["t2v", "i2v", "flf2v"], "positive_text_node": "7", "conditioning_node": "9", "guider_node": "29", "duration_head_file": "ltx-2.5-duration-head-bf16.safetensors", "min_seconds": 1.0, "length_targets": [{"node": "10", "input": "length"}, {"node": "27", "input": "frames_number"}, {"node": "50", "input": "length"}]}Models & dependencies#
Custom node packs (1)#
The non-core ComfyUI node classes this graph requires; the RunPod worker image the workflow runs on must bake or install a pack that provides every one of them.
ComfyUI-LTXVideoModels required (9)#
ltx-2.5-22b-dev-transformer-bf16.safetensorsgemma4-12b-with-proj-ltx-2.5-bf16.safetensorsltx-2.5-video-vae-bf16.safetensorsltx-2.5-audio-vae-bf16.safetensorsltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensorsltx-2.5-22b-distilled-lora-450-bf16.safetensorsltx-2.5-22b-ic-lora-ingredients-0.9.safetensorsltx-2.5-22b-ic-lora-pixel-spatial-upscaler-x2-1.0.safetensorsltx-2.5-duration-head-bf16.safetensorsOutput contract#
What a successful run of this workflow returns.
primary{"type": "video", "format": "mp4", "codec": "h264", "fps_source": "declared", "audio": true, "alpha": false, "description": "Clip with generated audio muxed in."}Taxonomy & routing#
How the control plane classifies this workflow — from the committed workflow-taxonomy-registry.json. It drives the consistency / control surface the agentic director can exercise over the workflow.
assetFamilyjoint-av-sceneoutputPackageProfilevideo-master-profilecontrolModalitiesmodel-lockprompt-template-locksampler-scheduler-lockseed-lockkeyframe-anchor-locktemporal-lockreference-ensembleconsistencyDimensionsidentitymotionenvironmentnotesLTX-2.5 audio+video (A.01.02, A.01.03). Each stage samples ONE concatenated audio+video latent. keyframe-anchor-lock applies in i2v (LTXVImgToVideoInplace in both stages) and flf2v (LTXVAddGuide at the first and last frame); reference-ensemble and the identity dimension apply in ingredients mode, where one reference sheet of characters, props and location conditions the clip through the Ingredients IC-LoRA.