Fish Audio upgrades its ASR model: speaker identification and inline emotion cues such as [laughter]

Fish Audio says its ASR model now ends boring transcripts. The upgraded speech-to-text system identifies different speakers within the same recording, understands how they feel, and writes emotion cues such as [laughter] and [surprised] directly into the transcript as labels, so subtitles, meeting notes and dialogue logs keep tone and non-verbal information instead of only getting the words right. The company states the model supports 83 languages and claims it is the most accurate STT model available, with access open now. For embodied AI and spoken human-robot interaction this is an upstream signal: speaker separation answers who is talking, emotion labels answer in what state, and both feed turn-taking and intent understanding for conversational agents operating in noisy real-world settings.





