Open joint audio-video generation on LTX-2.3 22B distilled (Apache 2.0): text-to-video or image-to-video with synchronized generated audio in ONE sampling pass (nested AV latent), 8-step distilled schedule at cfg 1 - the open audio-video flagship tier
Tags: flagshipltx-2videoaudiojoint-av2026-sota
Inputs (8)#
The typed parameter surface callers bind when they request this workflow. Enum options and numeric bounds are the values the workflow document declares.
modeenumdefault text_to_videotext_to_videoimage_to_videopromptstringdefault gentle rain over a temple garden, distant wind chimes, soft ambient tonesstart_imageimagestart_strengthfloatdefault 1.0min 0.0max 1.0generate_audiobooleandefault trueresolutionenumdefault landscape_1280landscape_1280portrait_1280base_768duration_secondsintegerdefault 6min 1max 20seedintegerdefault -1ComfyUI node graph (22)#
The executable ComfyUI prompt graph: 22 nodes across 21 distinct node classes, wired by 27 data dependencies. Nodes tinted green come from a custom node pack this workflow declares; the rest are ComfyUI core / baked-community classes.
Nodes (22)#
1CheckpointLoaderSimplecoreckpt_name = ltx-2.3-22b-distilled-fp8.safetensorsMODELCLIPVAE2LTXAVTextEncoderLoadercoretext_encoder = gemma_3_12B_it_fp4_mixed.safetensorsckpt_name = ltx-2.3-22b-distilled-fp8.safetensorsdevice = defaultCLIP3CLIPTextEncodecoretext = {{prompt}} tmplclip = ◂ node 2 · out[0]CONDITIONING4CLIPTextEncodecoretext = static image, jitter, warped motion, watermark, subtitles, distorted faces, harsh noise, clipping audioclip = ◂ node 2 · out[0]CONDITIONING5LTXVConditioningcorepositive = ◂ node 3 · out[0]negative = ◂ node 4 · out[0]frame_rate = 24.0CONDITIONINGCONDITIONING6EmptyLTXVLatentVideocorewidth = {{resolution_map[resolution].width}} tmplheight = {{resolution_map[resolution].height}} tmpllength = {{24 * duration_seconds + 1}} tmplbatch_size = 1LATENTstart_loadLoadImagecoreimage = {{start_image}} tmplIMAGEMASKstart_prepLTXVPreprocesscoreimage = ◂ node start_load · out[0]img_compression = 18IMAGEstart_applyLTXVImgToVideoInplacecorevae = ◂ node 1 · out[2]image = ◂ node start_prep · out[0]latent = ◂ node 6 · out[0]strength = {{start_strength}} tmplbypass = falseLATENTaudio_vaeLTXVAudioVAELoadercoreckpt_name = ltx-2.3-22b-distilled-fp8.safetensorsVAEaudio_latentLTXVEmptyLatentAudiocoreframes_number = {{24 * duration_seconds + 1}} tmplframe_rate = 24batch_size = 1audio_vae = ◂ node audio_vae · out[0]LATENTav_concatLTXVConcatAVLatentcorevideo_latent = ◂ node start_apply · out[0]audio_latent = ◂ node audio_latent · out[0]LATENT7RandomNoisecorenoise_seed = {{seed}} tmplNOISE8CFGGuidercoremodel = ◂ node 1 · out[0]positive = ◂ node 5 · out[0]negative = ◂ node 5 · out[1]cfg = 1.0GUIDER9KSamplerSelectcoresampler_name = euler_cfg_ppSAMPLER10ManualSigmascoresigmas = 1.0, 0.99375, 0.9875, 0.98125, 0.975, 0.909375, 0.725, 0.421875, 0.0SIGMAS11SamplerCustomAdvancedcorenoise = ◂ node 7 · out[0]guider = ◂ node 8 · out[0]sampler = ◂ node 9 · out[0]sigmas = ◂ node 10 · out[0]latent_image = ◂ node av_concat · out[0]LATENTLATENTav_splitLTXVSeparateAVLatentcoreav_latent = ◂ node 11 · out[0]LATENTLATENT12VAEDecodecoresamples = ◂ node av_split · out[0]vae = ◂ node 1 · out[2]IMAGEaudio_decodeLTXVAudioVAEDecodecoresamples = ◂ node av_split · out[1]audio_vae = ◂ node audio_vae · out[0]AUDIO13CreateVideocoreimages = ◂ node 12 · out[0]fps = 24audio = ◂ node audio_decode · out[0]VIDEO14SaveVideocorevideo = ◂ node 13 · out[0]filename_prefix = ltx2_av_sceneformat = autocodec = autoParameter banks (1)#
The prompt / configuration lookup tables this workflow keys into from its inputs — the vocabulary that turns a style / palette / preset selection into graph parameters.
resolution_map (3)#
landscape_1280{"width": 1280, "height": 704}portrait_1280{"width": 704, "height": 1280}base_768{"width": 768, "height": 512}Models & dependencies#
Models required (2)#
gemma_3_12B_it_fp4_mixed.safetensorsltx-2.3-22b-distilled-fp8.safetensorsOutput contract#
What a successful run of this workflow returns.
typevideoformatmp4audiotrueTaxonomy & routing#
How the control plane classifies this workflow — from the committed workflow-taxonomy-registry.json. It drives the consistency / control surface the agentic director can exercise over the workflow.
assetFamilyjoint-av-sceneoutputPackageProfilevideo-master-profilecontrolModalitiesmodel-locksampler-scheduler-lockseed-locktemporal-lockconsistencyDimensionsmotioncolor-scriptnotesLTX-2.3 22B distilled joint audio-video: one nested AV latent, one sampler, synchronized soundtrack; 8-step official distilled sigmas at cfg 1.