<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Asr</title><link>https://www.communeify.com/en/tag/asr/</link><description>Communeify - Your Community Platform</description><generator>Hugo</generator><atom:link href="https://www.communeify.com/en/tag/asr/" rel="self" type="application/rss+xml"/><item><title>AI Daily: Cohere-transcribe Open Source Speech Recognition - 2B Parameters, 3x Inference Efficiency, Top Choice for Enterprise Deployment</title><link>https://www.communeify.com/en/blog/cohere-transcribe-2b-open-source-asr-enterprise-guide/</link><guid isPermaLink="true">https://www.communeify.com/en/blog/cohere-transcribe-2b-open-source-asr-enterprise-guide/</guid><pubDate>Fri, 27 Mar 2026 08:54:46 +0800</pubDate><description>Built for Enterprise Production! How Cohere-transcribe Achieves 3x Inference Efficiency with 2B Parameters Does processing large amounts of audio data leave your server bills sky-high? Many of us have faced the dilemma: high accuracy usually comes with high computational costs. To be honest, this is a headache technical managers deal with daily.
Recently, Cohere released their first speech model, cohere-transcribe-03-2026. This is a speech-to-text model with 2B (2 billion) parameters, open-sourced under the business-friendly Apache 2.0 license. Trained from scratch on 14 key enterprise languages—including English, Chinese, Japanese, French, and German—it is specifically tailored for production environments and extreme efficiency.</description></item><item><title>Mistral Voxtral 4B Arrives: An Open-Source Real-Time Voice Model Under 500ms, Challenging Gemini and GPT-4o Dominance</title><link>https://www.communeify.com/en/blog/mistral-voxtral-4b-real-time-voice-ai-vs-gpt-4o-gemini/</link><guid isPermaLink="true">https://www.communeify.com/en/blog/mistral-voxtral-4b-real-time-voice-ai-vs-gpt-4o-gemini/</guid><pubDate>Thu, 05 Feb 2026 08:54:46 +0800</pubDate><description> This brand-new voice model not only boasts a compact 4-billion-parameter size but also breaks the rules of the voice transcription market with its stunning low latency and Apache 2.0 open-source license, bringing unprecedented local computing potential to developers.
In the past, when high-precision voice transcription was mentioned, people usually thought of OpenAI&amp;amp;rsquo;s Whisper or Google&amp;amp;rsquo;s voice services. While powerful, these tools often come with an annoying problem: latency. Typically, the system needs to wait for a sentence to finish, &amp;amp;ldquo;think&amp;amp;rdquo; for a moment, and then the text appears. For those wanting to build real-time interpretation or an AI assistant like Iron Man&amp;amp;rsquo;s Jarvis that can interrupt at any time, this wait is a fatal flaw.</description></item><item><title>Qwen3-ASR Heavyweight Open Source: Challenging Whisper&amp;#39;s Dominance, Precise Recognition for &amp;#39;Singing&amp;#39; and &amp;#39;Dialects&amp;#39;?</title><link>https://www.communeify.com/en/blog/ai-daily-qwen3-asr-opensource-vs-whisper-singing-dialect-recognition/</link><guid isPermaLink="true">https://www.communeify.com/en/blog/ai-daily-qwen3-asr-opensource-vs-whisper-singing-dialect-recognition/</guid><pubDate>Fri, 30 Jan 2026 08:54:46 +0800</pubDate><description>For a long time, OpenAI&amp;amp;rsquo;s Whisper series models have almost become the standard answer in the field of open source automatic speech recognition (ASR). Whenever developers need to handle speech-to-text tasks, the first name that comes to mind is usually it. But frankly, this &amp;amp;ldquo;one-player domination&amp;amp;rdquo; seems to be breaking. The Qwen team recently released the Qwen3-ASR series without warning. This is not just a routine version update, but more like a powerful impact on the boundaries of existing speech recognition technology.</description></item><item><title>Say Goodbye to Chopped Audio! Microsoft VibeVoice ASR Challenges 60-Minute Continuous Precise Transcription</title><link>https://www.communeify.com/en/blog/microsoft-vibevoice-asr-60-minute-long-form-transcription/</link><guid isPermaLink="true">https://www.communeify.com/en/blog/microsoft-vibevoice-asr-60-minute-long-form-transcription/</guid><pubDate>Thu, 22 Jan 2026 08:54:46 +0800</pubDate><description>Say Goodbye to Chopped Audio! Microsoft VibeVoice ASR Challenges 60-Minute Continuous Precise Transcription If you&amp;amp;rsquo;ve ever tried using AI to process long meeting minutes or podcast transcripts, the situation might feel familiar: the first ten minutes are accurate, but as the conversation gets longer, the semantics start to fall apart, or it even mixes up who said what.
This isn&amp;amp;rsquo;t because AI got stupider; the problem usually lies in &amp;amp;ldquo;segmentation&amp;amp;rdquo;.</description></item><item><title>MOSS-Transcribe-Diarize Released: Can this Multimodal AI Finally Understand Multi-person Arguments and Dialect Jokes?</title><link>https://www.communeify.com/en/blog/moss-transcribe-diarize-end-to-end-sats-multi-speaker/</link><guid isPermaLink="true">https://www.communeify.com/en/blog/moss-transcribe-diarize-end-to-end-sats-multi-speaker/</guid><pubDate>Fri, 09 Jan 2026 08:54:46 +0800</pubDate><description> OpenMOSS team released MOSS-Transcribe-Diarize at the beginning of 2026, an end-to-end multimodal large language model. It not only performs accurate speech transcription but also solves the long-standing problems of &amp;amp;ldquo;multi-person overlapping dialogue&amp;amp;rdquo; and &amp;amp;ldquo;emotional speech&amp;amp;rdquo; recognition. This article takes you deep into how this technology surpasses GPT-4o and Gemini and its practical application in complex speech scenarios.
(This article is a reserved post and will be updated later)
Have you ever had this experience? When reviewing video conference recordings or organizing interview audio, once two or three people speak at the same time, the subtitle software starts &amp;amp;ldquo;speaking gibberish,&amp;amp;rdquo; producing a pile of unintelligible text. Even when the speaker uses some dialect or gets emotional, AI often just waves the white flag.</description></item><item><title>Open Source ASR Newcomer GLM-ASR-Nano-2512 Debuts, Benchmarks Beat OpenAI Whisper V3</title><link>https://www.communeify.com/en/blog/glm-asr-nano-2512-opensource-asr-beats-openai-whisper-v3/</link><guid isPermaLink="true">https://www.communeify.com/en/blog/glm-asr-nano-2512-opensource-asr-beats-openai-whisper-v3/</guid><pubDate>Thu, 11 Dec 2025 08:54:46 +0800</pubDate><description> GLM-ASR-Nano-2512, with its lightweight design of 1.5B parameters, has beaten OpenAI Whisper V3 in multiple speech recognition benchmarks. This open-source model not only excels in dialect recognition such as Cantonese, but also accurately captures low-volume &amp;amp;ldquo;whisper&amp;amp;rdquo; conversations, providing developers and researchers with an efficient and powerful new choice.
In the field of Automatic Speech Recognition (ASR), OpenAI&amp;amp;rsquo;s Whisper series has long been seen as an insurmountable wall. Many developers are accustomed to using it as the default solution. However, with the iteration of technology, more competitive challengers are beginning to appear in the market. Recently, an open-source model named GLM-ASR-Nano-2512 has attracted widespread attention. It does not blindly pursue a huge parameter scale, but with a volume of 1.5B parameters, it demonstrates amazing efficiency and accuracy in handling complex real-world scenarios.</description></item><item><title>Meta AI&amp;#39;s Bombshell: How Does Omnilingual ASR Make 1,600 Languages &amp;#39;Speak&amp;#39;?</title><link>https://www.communeify.com/en/blog/meta-ai-omnilingual-asr-1600-languages-speech-recognition/</link><guid isPermaLink="true">https://www.communeify.com/en/blog/meta-ai-omnilingual-asr-1600-languages-speech-recognition/</guid><pubDate>Tue, 11 Nov 2025 08:54:46 +0800</pubDate><description> Meta AI has announced the revolutionary Omnilingual ASR technology, supporting speech recognition for over 1,600 languages, especially those with scarce resources. This open-source technology not only breaks technical bottlenecks but also hopes to truly bridge the language divide in the digital world through community power.
Have you ever thought about it? There are over 7,000 languages in the world, but on the internet, we mainly use only a few. This means that the native languages of billions of people are almost &amp;amp;lsquo;invisible&amp;amp;rsquo; in the digital world. This is not just a communication barrier, but a profound digital divide.</description></item><item><title>Alibaba&amp;#39;s Qwen Family Adds a Powerhouse! Introducing Qwen3-ASR-Flash, a New Way to Play with Speech Recognition?</title><link>https://www.communeify.com/en/blog/ali-qwen3-asr-flash-speech-recognition-revolution/</link><guid isPermaLink="true">https://www.communeify.com/en/blog/ali-qwen3-asr-flash-speech-recognition-revolution/</guid><pubDate>Tue, 09 Sep 2025 08:54:46 +0800</pubDate><description> Explore Alibaba&amp;amp;rsquo;s latest Qwen3-ASR-Flash speech recognition model. It not only supports 11 languages but also automatically detects language, filters noise, and achieves unimaginable accuracy. This article delves into its powerful features and practical applications, showing how this new AI star is changing the way we communicate.
Have you ever had this experience? You&amp;amp;rsquo;re in an important online meeting or listening to a high-value course, and you try to use a speech-to-text tool to take notes. The result is a garbled mess of text full of errors and nonsensical phrases, and you end up spending more time cleaning up the notes than you did in the meeting. This frustrating scenario is likely a shared memory for many.</description></item><item><title>Parakeet-TDT-0.6b-v3: NVIDIA&amp;#39;s New Open-Source Tool to Revolutionize Multilingual Speech-to-Text Experience</title><link>https://www.communeify.com/en/blog/parakeet-tdt-0-6b-v3-nvidia-opensource-multilingual-stt/</link><guid isPermaLink="true">https://www.communeify.com/en/blog/parakeet-tdt-0-6b-v3-nvidia-opensource-multilingual-stt/</guid><pubDate>Mon, 18 Aug 2025 08:54:46 +0800</pubDate><description> Explore NVIDIA&amp;amp;rsquo;s latest Parakeet-TDT-0.6b-v3 model, and how this 600-million-parameter AI model supports real-time speech-to-text for 25 European languages with amazing efficiency and accuracy, bringing new possibilities for developers and enterprises.
Have you ever wondered what it would be like if machines could effortlessly understand and record every word we say, whether in English, French, or Czech? It might sound like something out of a science fiction novel, but with the rapid development of artificial intelligence, this is no longer a distant dream.</description></item><item><title>Kyutai STT: Faster Than Whisper? A French AI Challenger Pushes the Limits of Real-Time Speech Recognition</title><link>https://www.communeify.com/en/blog/kyutai-stt-faster-than-whisper-realtime-speech-recognition/</link><guid isPermaLink="true">https://www.communeify.com/en/blog/kyutai-stt-faster-than-whisper-realtime-speech-recognition/</guid><pubDate>Mon, 23 Jun 2025 10:07:46 +0800</pubDate><description> Discover Kyutai STT, the open-source speech-to-text model from France that challenges OpenAI Whisper in both speed and accuracy, built specifically for real-time interaction. Whether you’re a developer, researcher, or AI enthusiast, this article explores what sets it apart.
You might be thinking, &amp;amp;ldquo;Another speech recognition model? Aren’t there already enough options?&amp;amp;rdquo;
Honestly, I thought the same at first. But after diving into Kyutai STT (Speech-to-Text)—the latest open-source release from the French AI lab Kyutai—I realized it’s something different. This isn’t just another transcription tool. It’s purpose-built for real-time interaction, and it achieves an impressive balance between latency and accuracy.</description></item><item><title>NVIDIA Parakeet Speech Recognition Model: 600M Parameters to Challenge OpenAI? Transcribe 60-Minute Audio in 1 Second, Open-Source and Powerful!</title><link>https://www.communeify.com/en/blog/nvidia-parakeet-speech-model-opensource/</link><guid isPermaLink="true">https://www.communeify.com/en/blog/nvidia-parakeet-speech-model-opensource/</guid><pubDate>Thu, 08 May 2025 08:00:46 +0800</pubDate><description> The field of AI speech recognition is surging! NVIDIA&amp;amp;rsquo;s recently open-sourced Parakeet TDT 0.6B V2 model on Hugging Face has quickly become a focal point with its amazing transcription speed, accuracy comparable to commercial tools, and generous open-source license. What magical power does this &amp;amp;rsquo;little parakeet&amp;amp;rsquo; possess? Let&amp;amp;rsquo;s take a look!
The field of AI speech recognition has been bustling with activity recently! Major tech giants are all gearing up in this race, constantly releasing more powerful models. And not long ago, the graphics chip leader NVIDIA also dropped a bombshell—they open-sourced a model called nvidia/parakeet-tdt-0.6b-v2 on the well-known AI community platform Hugging Face. This is not just some new toy; it&amp;amp;rsquo;s a secret weapon specifically designed for high-quality English automatic speech recognition (ASR) and dictation.</description></item></channel></rss>