Skip to main content
LTXV Reference Audio transfers the voice identity of a speaker from a reference audio clip to generated audio. It encodes the reference audio into the conditioning and optionally patches the model with identity guidance, which runs an extra forward pass without the reference each step to amplify the speaker identity effect.

Inputs

Note: Identity guidance is only applied when identity_guidance_scale is greater than 0 and the current sampling step is within the range defined by start_percent and end_percent. The reference audio is resampled to the audio VAE’s sample rate if the two differ.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): ae15c5838656324667d099614b325b863341f05afda43054658999574522dd49