Overview
ComfyUI Partner Nodes give you access to audio generation models for music, speech, and sound effects.Browse audio models
Seed Audio 1.0
Text-to-audio and text-to-music, including instrumental and vocal tracks
Sonilo
Generate a full soundtrack that matches a video’s rhythm and content
What you can do with audio models
- Text-to-audio: Generate music or sound effects from a text description
- Video-to-music: Create a soundtrack that matches a video’s content and pacing
- Voice & speech: Generate speech or singing from text and reference audio
Choosing a model by task
Music & soundtracks
- Seed Audio 1.0 (ByteDance) — text-to-audio and text-to-music generation, including instrumental and vocal tracks
- Sonilo (Sonilo) — generate a full soundtrack from a video: it analyzes the video’s rhythm and content to produce a matching score
Sound effects & utility
- Seed Audio 1.0 (ByteDance) — also supports sound effects and audio-style transfer
All audio partners
See Pricing for per-model rates, and Concurrency Limits for usage caps.