{"id":23527,"date":"2026-08-26T17:06:02","date_gmt":"2026-08-26T17:06:02","guid":{"rendered":"https:\/\/scannn.com\/introducing-gemini-3-5-transcribe\/"},"modified":"2026-08-26T17:06:02","modified_gmt":"2026-08-26T17:06:02","slug":"introducing-gemini-3-5-transcribe","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/introducing-gemini-3-5-transcribe\/","title":{"rendered":"Introducing Gemini 3.5 Transcribe"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p data-block-key=\"jlfw0\">Today, we\u2019re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.<\/p>\n<p data-block-key=\"3k0f9\">Across our products like the Gemini app and on Android, we\u2019ve seen consumers already benefiting from this transcription model with new voice capabilities like Rambler on Android and in the Gemini app on macOS. Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.<\/p>\n<p data-block-key=\"1vu9n\">We&#8217;ve built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you\u2019re building voice agents, real-time captioning tools, or post-call analytics pipelines. The model is available across two separate APIs:<\/p>\n<ul>\n<li data-block-key=\"9eqe4\"><b>Real-time streaming:<\/b> Delivers continuous, bidirectional streaming with sub-second latency for interactive voice apps via the Live API using <b><code>gemini-3.5-transcribe-live<\/code><\/b><b>.<\/b><\/li>\n<li data-block-key=\"dlsm9\"><b>Pre-recorded audio processing:<\/b> Transcribes recorded audio, meetings, call logs, and more with speaker attribution and word-level timestamps via the Interactions API using <b><code>gemini-3.5-transcribe<\/code><\/b><b>.<\/b><\/li>\n<\/ul>\n<h2 data-block-key=\"ddu74\">Get more precise and intelligent transcription<\/h2>\n<p data-block-key=\"8h861\">Gemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice.<\/p>\n<ul>\n<li data-block-key=\"cqo82\"><b>Smart transcription:<\/b> Seamlessly handles self-corrections (like <i>&#8220;let\u2019s meet Tuesday\u2014no, Wednesday&#8221;<\/i>), removes filler words (\u201cums\u201d and \u2018\u201cahs&#8221;), auto-formats your text.<\/li>\n<li data-block-key=\"2pddt\"><b>Function calling:<\/b> The model can delegate complex tasks (such as image generation and file analysis) to other Gemini models via function calls. Currently available in the Gemini macOS app.<\/li>\n<li data-block-key=\"a6iq0\"><b>More precise transcription:<\/b> As measured by Artificial Analysis, achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases. It shows strong performance across noisy, real-world environments, accurately capturing alphanumeric entities like postal codes and order IDs.<\/li>\n<li data-block-key=\"2aqjp\"><b>Custom vocabulary:<\/b> Recognizes specialized jargon and unique spellings by seamlessly adapting transcriptions to your provided custom vocabulary.<\/li>\n<li data-block-key=\"b6935\"><b>Global language support:<\/b> Automatically detects and transcribes over 85 languages, seamlessly handling regional accents and diverse dialects.<\/li>\n<li data-block-key=\"1v87a\"><b>Multi-speaker identification:<\/b> Accurately attributes speech in pre-recorded audio with timestamps for up to three speakers (support for 3+ speakers is experimental).<\/li>\n<\/ul>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/blog.google\/innovation-and-ai\/models-and-research\/gemini-models\/gemini-3-5-transcribe\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Today, we\u2019re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text. Across our products like the Gemini app and on Android, we\u2019ve [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":23528,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[100],"tags":[],"class_list":["post-23527","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-google"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23527","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=23527"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23527\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/23528"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=23527"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=23527"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=23527"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}