One VACE graph with four jobs: replace a masked region, restyle a whole clip under a control signal, extend a clip past its last frame, or outpaint its canvas. Which job it does is decided by what the control video and mask track are fed, which is why the node set changes after rendering.
Tags: motionwan2.1vacevideo-to-videoinpaintoutpaintrunpod-serverlessvolume-backed
Inputs (22)#
The typed parameter surface callers bind when they request this workflow. Enum options and numeric bounds are the values the workflow document declares.
source_videovideorequiredmodeenumdefault restyleinpaintrestyleextendoutpaintcontrolenumdefault rawnonedepthposecannyrawmask_videovideoreference_imageimagekeep_framesintegerdefault 16min 1max 236pad_leftintegerdefault 0min 0max 512pad_rightintegerdefault 0min 0max 512pad_topintegerdefault 0min 0max 512pad_bottomintegerdefault 0min 0max 512promptstringrequireddefault the same scene, cinematic lightingnegative_promptstringdefault 色调艳丽,过曝,静态,细节模糊不清,字幕,风格,作品,画作,画面,静止,整体发灰,最差质量,低质量,JPEG压缩残留,丑陋的,残缺的,多余的手指,画得不好的手部,画得不好的脸部,畸形的,毁容的,形态畸形的肢体,手指融合,静止不动的画面,杂乱的背景,三条腿,背景人很多,倒着走strengthfloatdefault 1.0min 0.0max 2.0resolutionenumdefault 480p_landscape_832x480480p_landscape_832x480480p_portrait_480x832720p_landscape_1280x720720p_portrait_720x1280lengthintegerdefault 81min 5max 241fpsfloatdefault 16.0min 1.0max 30.0stepsintegerdefault 20min 1max 60cfgfloatdefault 5.0min 1.0max 20.0shiftfloatdefault 8.0min 0.0max 20.0seedintegerdefault 0min 0samplerenumdefault uni_pcuni_pceulerdpmpp_2mschedulerenumdefault simplesimplenormalbetakarrasComfyUI node graph (14)#
The executable ComfyUI prompt graph: 14 nodes across 13 distinct node classes, wired by 18 data dependencies. Nodes tinted green come from a custom node pack this workflow declares; the rest are ComfyUI core / baked-community classes.
Nodes (14)#
1UNETLoadercoreunet_name = wan2.1_vace_14B_fp16.safetensorsweight_dtype = defaultMODEL2CLIPLoadercoreclip_name = umt5_xxl_fp16.safetensorstype = wandevice = defaultCLIP3VAELoadercorevae_name = wan_2.1_vae.safetensorsVAE4ModelSamplingSD3coremodel = ◂ node 1 · out[0]shift = {{shift}} tmplMODEL5CLIPTextEncodecoreclip = ◂ node 2 · out[0]text = {{prompt}} tmplCONDITIONING6CLIPTextEncodecoreclip = ◂ node 2 · out[0]text = {{negative_prompt}} tmplCONDITIONING7LoadVideocorefile = {{source_video}} tmplVIDEO8GetVideoComponentscorevideo = ◂ node 7 · out[0]IMAGEAUDIOFLOATCOMBOCOMBO9WanVaceToVideocorepositive = ◂ node 5 · out[0]negative = ◂ node 6 · out[0]vae = ◂ node 3 · out[0]width = {{resolution_map[resolution].width}} tmplheight = {{resolution_map[resolution].height}} tmpllength = {{length}} tmplbatch_size = 1strength = {{strength}} tmplcontrol_video = ◂ node 8 · out[0]CONDITIONINGCONDITIONINGLATENTINT10KSamplercoremodel = ◂ node 4 · out[0]positive = ◂ node 9 · out[0]negative = ◂ node 9 · out[1]latent_image = ◂ node 9 · out[2]seed = {{seed}} tmplsteps = {{steps}} tmplcfg = {{cfg}} tmplsampler_name = {{sampler}} tmplscheduler = {{scheduler}} tmpldenoise = 1.0LATENT11TrimVideoLatentcoresamples = ◂ node 10 · out[0]trim_amount = ◂ node 9 · out[3]LATENT12VAEDecodecoresamples = ◂ node 11 · out[0]vae = ◂ node 3 · out[0]IMAGE13CreateVideocoreimages = ◂ node 12 · out[0]fps = {{fps}} tmplVIDEO14SaveVideocorevideo = ◂ node 13 · out[0]filename_prefix = wan21-vaceformat = mp4Parameter banks (3)#
The prompt / configuration lookup tables this workflow keys into from its inputs — the vocabulary that turns a style / palette / preset selection into graph parameters.
resolution_map (4)#
480p_landscape_832x480{"width": 832, "height": 480}480p_portrait_480x832{"width": 480, "height": 832}720p_landscape_1280x720{"width": 1280, "height": 720}720p_portrait_720x1280{"width": 720, "height": 1280}post_render (1)#
vace_modes{"mode_input": "mode", "control_input": "control", "mask_video_input": "mask_video", "reference_image_input": "reference_image", "keep_frames_input": "keep_frames", "pad_inputs": {"left": "pad_left", "right": "pad_right", "top": "pad_top", "bottom": "pad_bottom"}, "target_class": "WanVaceToVideo", "frames_class": "GetVideoComponents", "frames_output": 0}requires_families (3)#
wan21-vacewan-sharedcontrolnet-auxModels & dependencies#
Custom node packs (1)#
The non-core ComfyUI node classes this graph requires; the RunPod worker image the workflow runs on must bake or install a pack that provides every one of them.
comfyui_controlnet_auxModels required (3)#
wan2.1_vace_14B_fp16.safetensorsumt5_xxl_fp16.safetensorswan_2.1_vae.safetensorsOutput contract#
What a successful run of this workflow returns.
primary{"type": "video", "format": "mp4", "codec": "h264", "fps_source": "declared", "audio": false, "alpha": false, "description": "The edited clip."}Taxonomy & routing#
How the control plane classifies this workflow — from the committed workflow-taxonomy-registry.json. It drives the consistency / control surface the agentic director can exercise over the workflow.
assetFamilyregion-edited-videooutputPackageProfilevideo-master-profilecontrolModalitiesmodel-lockprompt-template-locksampler-scheduler-lockseed-locktemporal-lockcontrolnet-depthcontrolnet-posecontrolnet-cannycontrolnet-inpaintreference-ensembleconsistencyDimensionsidentitymotionlightingenvironmentnotesWan 2.1 VACE, four modes in one graph. Every controlnet-* modality is listed because the `control` input really does select between depth, pose and canny preprocessors, and `inpaint` mode really does take a mask track; reference-ensemble covers the optional reference still. The mode is a node-set change applied after rendering, not a value substitution.