Skip to main content

Overview

ComfyUI Partner Nodes give you access to audio generation models for music, speech, and sound effects.

Browse audio models

Seed Audio 1.0

Text-to-audio and text-to-music, including instrumental and vocal tracks

Sonilo

Generate a full soundtrack that matches a video’s rhythm and content

What you can do with audio models

  • Text-to-audio: Generate music or sound effects from a text description
  • Video-to-music: Create a soundtrack that matches a video’s content and pacing
  • Voice & speech: Generate speech or singing from text and reference audio

Choosing a model by task

Music & soundtracks

  • Seed Audio 1.0 (ByteDance) — text-to-audio and text-to-music generation, including instrumental and vocal tracks
  • Sonilo (Sonilo) — generate a full soundtrack from a video: it analyzes the video’s rhythm and content to produce a matching score

Sound effects & utility

  • Seed Audio 1.0 (ByteDance) — also supports sound effects and audio-style transfer

All audio partners

See Pricing for per-model rates, and Concurrency Limits for usage caps.