Google: Gemini 3.1 Flash TTS Preview

Gemini 3.1 Flash TTS Preview is Google's text-to-speech model, released 2026-04, a substantial step up from Gemini 2.5 Flash TTS. It covers 70+ languages and introduces 200+ inline audio tags for steering delivery, emotion, and pacing mid-sentence. The model supports two speakers with independent voice and style configuration, outputs 24 kHz / 16-bit mono PCM, and watermarks output with SynthID.

Modalities
inTextoutAudio
Price / 1M−5%
$0.95 / $19.00$1.00 / $20.00
Context
33K
Knowledge cutoff

Capabilities

Structured output

Channels

Same weights on every channel. Requests route to the cheapest healthy one; the channel used is printed on each billing line.

Provider
Discount
Input /M
Output /M
Cache Read /M
Cache Create 5m /M
Cache Create 1h /M
Defaultstable
−5%
$0.95
$19.00

Quick Start

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.mirapi.ai/v1",   # changed
    api_key=os.environ["MIRAPI_API_KEY"],   # changed
)

resp = client.chat.completions.create(
    model="google/gemini-3.1-flash-tts-preview",
    messages=[{"role": "user", "content": "Hello"}],
)
Wire Gemini 3.1 Flash TTS Preview into your product

1M free tokens to start · no card · OpenAI-compatible, one base_url change

Start free