Google DeepMind·· 2026-08-27精选AI 评分71
Google DeepMind 发布 Gemini 3.5 Transcribe 语音转写模型
Intelligent transcription with Gemini 3.5 Transcribe
AI 导读
Google DeepMind 发布 Gemini 3.5 Transcribe,称其为目前最精确的语音转文字模型,可直接把原始音频转成准确、格式化后的文本。据 Artificial Analysis 测量,流式场景平均词错误率为 4.0%,非流式为 2.6%,最终转写时间较前代 Chirp 3 改善 70%。
推荐理由
官方给出流式与非流式 WER 及延迟对比,可据此判断语音转写接入现有工作流的门槛。
来源:Google DeepMind · deepmind.google