#Audio pitch mismatch between playground file vs WebSocket API file

1 messages · Page 1 of 1 (latest)

stable glade
#

Hey team,
We encountered an issue while trying to download some lengthy audio files (~2k words) using the playground. Looks like playground does not support large "transcript" size. So, we opted for WebSocket API to convert our text to speech.
Previously, we downloaded an audio file via the playground with the Voice: Sweet Lady and the speed set to -0.2. However, when we tried to generate the same audio using the API with the same parameters, the result was noticeably different—the voice seemed higher pitched compared to the one we got from the playground.
Could you help us figure out:
How we can get similar audio outputs using the API?
Is there a way to download larger audio files on the playground without it hanging?
Sharing audios here for your reference (first one is from playground , second one is from Websocket API)

This is the payload I'm sending

{"transcript":"xxxxx","voice":{"mode":"id","id":"e3827ec5-697a-4b7c-9704-1a23041bbc51", "__experimental_controls": {"speed": -0.2}},"output_format":{"container":"raw","encoding":"pcm_f32le","sample_rate":44100},"model_id":"sonic-english"}

Thanks!

stable glade
#

@vague wren @crude acorn can you please help here? 🙏

raven mantle
#

Just a Cartesia user here, but you are talking pitch and then in your payload, you are modifying speed.

#

Duration of both clips is quite similar as well.

#

Did you try a new generation?

#

Each generation will have its own variations in tone and pronunciation.