Tags: audiovoicespeechttsnarration
Inputs (4)#
The typed parameter surface callers bind when they request this workflow. Enum options and numeric bounds are the values the workflow document declares.
scriptstringrequiredvoice_profileenumdefault af_bellaaf_bellaaf_nicoleaf_saraham_adamam_michaelbf_emmabf_isabellabm_georgebm_lewisspeaking_ratefloatdefault 1.0min 0.7max 1.3seedintegerdefault -1ComfyUI node graph (4)#
The executable ComfyUI prompt graph: 4 nodes across 4 distinct node classes, wired by 3 data dependencies. Nodes tinted green come from a custom node pack this workflow declares; the rest are ComfyUI core / baked-community classes.
Nodes (4)#
1KokoroTextToSpeechcustom packtext = {{script}} tmplvoice = {{voice_profile}} tmplspeed = {{speaking_rate}} tmplAUDIO2AudioDeessercustom packaudio = ◂ node 1 · out[0]strength = 0.3AUDIO3AudioLoudnessNormalizecustom packaudio = ◂ node 2 · out[0]target_lufs = -16.0AUDIO4SaveAudiocoreaudio = ◂ node 3 · out[0]filename_prefix = voice_performance_synthesisPrompt construction#
template{script} | voice profile: {voice_profile}, speech rate: {speaking_rate}Variables (3)#
script{{script}} tmplvoice_profile{{voice_profile}} tmplspeaking_rate{{speaking_rate}} tmplModels & dependencies#
Custom node packs (3)#
The non-core ComfyUI node classes this graph requires; the RunPod worker image the workflow runs on must bake or install a pack that provides every one of them.
AudioDeesserAudioLoudnessNormalizeKokoroTextToSpeechModels required (2)#
kokoro-v1.0.onnxkokoro-voices-v1.0.binOutput contract#
What a successful run of this workflow returns.
typeaudioformatflaccodecpcm_s16lechannels1descriptionNarration/voice performance masteraudio_package{"mixdown": {"artifact_id": "audio_mixdown", "path_template": "audio/{{job_id}}/mixdown/voice-master.wav", "format": "wav", "sample_rate_hz": 48000, "bit_depth": 24, "channels": 1, "codec": "pcm_s24le"}, "stems": {"format": "wav", "sample_rate_hz": 48000, "bit_depth": 24, "channels": 1, "artifacts": [{"stem": "dialogue", "artifact_id": "stem_dialogue", "path_template": "audio/{{job_id}}/stems/dialogue.wav", "required": true}, {"stem": "ambience", "artifact_id": "stem_room_tone", "path_template": "audio/{{job_id}}/stems/room-tone.wav", "required": false}], "archive": {"artifact_id": "stems_archive", "path_template": "audio/{{job_id}}/stems/stems.zip", "format": "zip"}}, "multi_track": {"format": "wav", "sample_rate_hz": 48000, "bit_depth": 24, "channels": 2, "artifacts": [{"track": "dialogue", "artifact_id": "multitrack_dialogue", "path_template": "audio/{{job_id}}/multitrack/dialogue.wav", "required": true}, {"track": "sfx", "artifact_id": "multitrack_sfx", "path_template": "audio/{{job_id}}/multitrack/sfx.wav", "required": false}, {"track": "music", "artifact_id": "multitrack_music", "path_template": "audio/{{job_id}}/multitrack/music.wav", "required": false}], "archive": {"artifact_id": "multitrack_archive", "path_template": "audio/{{job_id}}/multitrack/multitrack.zip", "format": "zip"}}, "loudness": {"artifact_id": "loudness_report", "path_template": "audio/{{job_id}}/analysis/loudness-report.json", "format": "json", "standard": "broadcast", "target_integrated_lufs": -18, "max_true_peak_dbtp": -2}, "compliance": {"artifact_id": "audio_compliance_report", "path_template": "audio/{{job_id}}/analysis/audio-compliance-report.json", "format": "json", "loudness_tolerance_lufs": 1.0, "true_peak_tolerance_dbtp": 0.3, "max_clipping_percent": 0.1, "enforce_format_normalization": true}, "cue_sheet": {"artifact_id": "cue_sheet", "path_template": "audio/{{job_id}}/metadata/cue-sheet.csv", "format": "csv", "include_timecode": true, "include_beat_markers": false, "include_sections": true}}Taxonomy & routing#
How the control plane classifies this workflow — from the committed workflow-taxonomy-registry.json. It drives the consistency / control surface the agentic director can exercise over the workflow.
assetFamilyanim-lipsync-viseme-sheetoutputPackageProfileaudio-master-profilecontrolModalitiescamera-lockcolor-script-lockdepth-constraintidentity-adapteridentity-locklatent-reusemodel-lockpalette-lockpose-constraintprompt-template-lockreference-ensemblesampler-scheduler-lockscene-lockseed-lockstyle-anchorstyle-lock +2 moreconsistencyDimensionsidentitymotioncolor-scriptnotesAuto-mapped from audio defaults