Hey guys!
We’ve been trying cartesia new STT model for a little while now and to be honest it is one if not the faster STT model we encountered, with high accuracy and very low latency !
the only encounter we meet is :
We have sometimes detections of « … » unknown speeches like hesitating for us ( in french) and it kind of screws our llm and tts after!
Wanted to know if any of you already met this issue and if so, did find a way to solve it ?
We are working on a rpi so audio processing is also a bit different