Google for Developers’ cover photo
Google for Developers

Google for Developers

Technology, Information and Internet

Mountain View, CA 4,302,389 followers

Join a community of creative developers and learn how to use the latest in technology—from AI and cloud, to mobile & web

About us

Discover the latest technologies, resources, events, and announcements to help you build smarter and ship faster. Explore more at developers.google.com

Website
http://developers.google.com
Industry
Technology, Information and Internet
Company size
10,001+ employees
Headquarters
Mountain View, CA
Specialties
coding, engineering, firebase, android, cloud, web development, and mobile development

Updates

  • Introducing Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. These new audio models let you build voice agents that reason and execute complex tasks in the background while maintaining natural-sounding conversations. To build a truly fluid voice agent, you need a model that can handle interruptions, process context in near real time, and execute tasks without breaking the conversational flow. That level of responsiveness and reliability is difficult when stitching together separate models in the traditional cascaded pipeline. These models replace that with a single, natively multimodal model—making it simpler and more cost-effective to build voice agents. Here’s how you can build with them: 🔊 Gemini 3.8 Live helps you build interactive, high-volume voice assistants like a live onboarding agent or customer support bot. Because it handles async tool calling and visual processing natively, your agent can fetch database records or trigger APIs in the background without creating awkward pauses in the conversation. 🔊 Gemini 3.8 Live Extended Thinking lets you create agents for complex, multi-step workflows like a live coding assistant. It reasons and speaks simultaneously, narrating its thought process in near real time using natural verbal cues, keeping users engaged while it reasons through heavy computational tasks. Available today via the Gemini API on Google AI Studio and coming soon to the Gemini Enterprise Agent Platform. Check out the blog to learn more: https://goo.gle/46YPGAd Watch the demo below to see Gemini 3.8 Extended Thinking in action as a workshop assistant. It guides and troubleshoots cyberdeck building in real time using live camera and voice. 

  • Build richer music experiences into your apps with Lyria 3.5, our cleanest, most controllable music generation model yet 🎵✨ With the Gemini API, developers can ship production-ready music featuring significantly reduced noise, distortion, and artifacts. Now devs have more granular generation control with authentic regional accents across multiple languages, strict lyric adherence, and native handling for bracketed cues like [guitar solo]. Try Lyria 3.5 in Google AI Studio or explore the developer documentation → https://goo.gle/4gMk7OC

  • Introducing Gemini 3.8 Flash ⚡️ Built to be your sharpest coding partner, 3.8 Flash brings major upgrades over 3.7 Flash across SWE and agentic tasks. You get dynamic thinking controls to balance reasoning depth with your budget— all at the same introductory price of $0.75/1M input and $3.75/1M output tokens through the end of the year. To see it in action, watch 3.8 Flash build this app from scratch in Antigravity. We gave it a single goal to create a dynamic UI with the Google Maps API and the model tackled the complex API integrations and wrote all of the code completely autonomously. Try it today via Google Antigravity and the Gemini API via Google AI Studio and Android Studio. More details in the blog: https://goo.gle/46B6Fs3

  • 📽️⚡ Process long-form video more with significantly fewer tokens without sacrificing quality. Agentic video understanding unlocks new ways to process long-form video to deliver massive token reductions (up to 88%), better accuracy, and lower costs. With this feature, Gemini actively scans visual frames, audio, and transcripts instead of passively ingesting video at 1 FPS. Now you can tackle complex video workflows like: ▸ Sub-second retrieval: Catch split-second cuts missed at 1 FPS.  ▸ Long-form search: Query multi-hour videos without overflowing context.  ▸ Anomaly detection: Resample interesting windows at higher FPS.  ▸ Counting: Accurately track repeated actions and objects over time. And more. Learn here: https://goo.gle/4gFY2kK Agentic video understanding is supported by Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite and is available today for video uploads and YouTube videos via the Gemini API on Google AI Studio and Gemini Enterprise Agent Platform. See how agentic video understanding improves token efficiency in Gemini 3.7 Flash when compared with static video processing:

  • Last week we introduced Gemini 3.5 Transcribe, our latest text-to-speech model. But, what does this actually mean for your projects? We built this app in Google AI Studio to demonstrate just how much smarter 3.5 Transcribe is. When streaming live audio simultaneously through two parallel transcription pipeline modes, you can see how smart transcription automatically strips out disfluencies (ums, ahs) and condenses long, rambling thoughts while verbatim keeps exact match transcription live:

  • View organization page for Google for Developers

    4,302,389 followers

    🗣️✨ Speak your vibes into existence with Gemini 3.5 Transcribe The latest audio model is built for apps that need live transcription and quick visual prototyping. Watch Geneviève Huskens turn spoken ideas into a dynamic mood board using Gemini 3.5 Transcribe on the Live API. The model maps speech to user intent to create prompts that trigger live image generation and editing with Nano Banana 2 Lite: Try the voice-powered app yourself: https://goo.gle/4wYBsdo

Affiliated pages

Similar pages