Create natural speech across supported locales and voices, then carry the audio into connected video workflows.