This page documents production updates to Cloud STT. You can periodically check this page for announcements about new or updated features, bug fixes, known issues, and deprecated functionality.
You can see the latest product updates for all of Google Cloud on the Google Cloud page, browse and filter all release notes in the Google Cloud console, or programmatically access release notes in BigQuery.
To get the latest product updates delivered to you, add the URL of this page to your feed reader, or add the feed URL directly.
November 13, 2025
Speech-to-Text has just launched chirp_3 in public preview for the regions asia-south1, europe-west2, europe-west3 and northamerica-northeast1.
For more information about the Chirp 3 model, see Chirp 3 Transcription: Enhanced multilingual accuracy.
October 13, 2025
Speech-to-Text is excited to announce the General Availability (GA) of the Chirp 3: Transcription, the latest generation of Google's multilingual Automatic Speech Recognition (ASR)-specific generative model, delivering state-of-the-art ASR accuracy and multilingual capabilities. Available exclusively in the Speech-to-Text API V2, Chirp 3 delivers significant enhancements in transcription accuracy and speed over previous versions.
Under the new chirp_3 model identifier, you can now leverage powerful new capabilities, including speaker diarization to identify different speakers and automatic language detection for multilingual audio. The model supports all major recognition methods —StreamingRecognize, Recognize, and BatchRecognize- making it suitable for both real-time and batch processing. Chirp 3 also offers advanced features such as speech adaptation for custom vocabularies and a built-in denoiser to improve results from noisy audio.
To explore the new Chirp 3: Transcription model's capabilities and learn how to leverage its full potential, please visit our updated documentation page.
August 29, 2025
Speech-to-Text has just launched chirp_3 in Public Preview. With this Public Preview, Chirp 3: Transcription we are now expanding language transcription in more than 85+ languages and locales, in addition to StreamingRecognize and SyncRecognize requests for real-time and short-form audio. Under the chirp_3 model flag, you can experience significant improvements in accuracy and speed, and leverage powerful features like Speaker Diarization and language-agnostic transcription.
To explore the new Chirp 3: Transcription model's capabilities and learn how to leverage its full potential, please visit our official documentation page.
April 11, 2025
Speech-to-Text has launched chirp_3 in Private Preview. Chirp 3: Transcription is the latest generation of Google's multilingual Automatic Speech Recognition (ASR)-specific generative models that further enhances its ASR accuracy and multilingual capabilities. Under the new chirp_3 model flag, you can experience significant improvements in accuracy and speed, and leverage powerful new features like Speaker Diarization and language-agnostic transcription. Chirp 3 supports BatchRecognize requests within the Speech-to-Text v2 API, making it ideal for transcribing long-form audio.
To explore the new Chirp 3: Transcription model's capabilities and to learn how to leverage its full potential, please visit our official documentation page. To gain access to this Private Preview, please contact our sales team.
January 27, 2025
Speech-to-Text is generally available (GA) in the Chirp 2 model in asia-southeast1, us-central1, and europe-west4.
For more information about the Chirp 2 model, see Chirp 2: Enhanced multilingual accuracy. For code samples, see Get started with Chirp 2 using Speech-to-Text V2 SDK in GitHub.
October 07, 2024
Speech-to-Text has updated the Generally Available Chirp 2 model, further enhancing its ASR accuracy and multilingual capabilities. Under the existing chirp_2 model flag, you can experience significant improvements in accuracy and speed, as well as support for word-level timestamps, model adaptation, and speech translation. Finally, Chirp 2 can support Streaming Recognizer requests, in addition to the already supported Sync and Batch Recognition requests, allowing its use in realtime applications.
Explore the new chirp_2 model's capabilities and learn how to leverage its full potential by visiting our updated documentation and tutorials.
January 09, 2024
Model adaptation is now available for latest_long models in 13 languages. Also, its quality was substantially improved for latest_short models. To determine whether this feature is available for your language, see Language support.
January 08, 2024
Speech-to-Text has launched a new model, named chirp_telephony to bring the accuracy gains of our chirp model to telephony-specific use cases. The new model is a fine-tuned version of our very successful chirp model, based on the Universal large Speech Model(USM) architecture, on audio that originated from a phone call typically recorded at an 8 kHz sampling rate. For more information, see Speech-to-Text supported languages.
November 06, 2023
Speech-to-Text has launched two models, named telephony and telephony_short. The two models are customized to recognize audio that originates from a phone call and corresponds to the most recent versions of the existing phone_call model. For more information, see Speech-to-Text supported languages.
February 07, 2023
We are removing SpeechContext.strength field within the next 4
weeks, because it has been deprecated and unused for more than a year. The
documentation doesn't have references to this field anymore, and the clients aren't supposed to use it.
November 11, 2022
Speech-to-Text has updated its pricing policy. Enhanced models are no longer priced differently than standard models. Usage of all models will be reported to and priced like standard models. Also, all Cloud Speech-to-Text requests will now be rounded up to the nearest 1 second, with no minimum audio length (requests were previously rounded up to the nearest 15 seconds). See the Pricing page for details.
October 03, 2022
Speaker Diarization is now available for "Latest" models in en-US. This feature recognizes multiple speakers in the same audio clip. Latest models use a new model for diarization from previous models. For more information see Speaker Diarization.
April 21, 2022
"Latest" models are available in more than 20 languages. These models employ new end-to-end machine learning techniques and can improve the accuracy of your recognized speech. For more information see Latest models.
November 08, 2021
Speech-to-Text has launched two new medical speech models, which are tailored for recognition of words that are common in medical settings. See the medical models documentation for more details.
July 21, 2021
Speech-to-Text has launched a GA version of the Spoken Emoji and Spoken Punctuation features. See the documentation for details.
June 28, 2021
The Speech-to-Text now supports multi-region endpoints as a GA feature. See the multi-region endpoints documentation for more information.