API & Backend

Real-Time AI Streaming with FastAPI and Google Gemini API (SSE)

Learn how to build low-latency Server-Sent Events (SSE) streaming endpoints in FastAPI using the official Google GenAI SDK and structured tool calling.

Updated: September 9, 20265 min
Share:XLinkedIn

Summary & Direct Solution (TL;DR)

Integrating FastAPI with Google Gemini 3.7 API using Server-Sent Events (SSE) delivers initial tokens to clients in under 250ms. By orchestrating Python asynchronous generators, HTTP/2 streaming, and Pydantic v2 schemas, engineering teams achieve high concurrency while eliminating perceived latency.

Key Technical Takeaways:

  • TTFT (Time to First Token) Optimization: Streaming yields immediate output within 200-300ms rather than waiting 5-10 seconds for complete generation.
  • FastAPI StreamingResponse & SSE Protocol: text/event-stream headers enable effortless consumption with browser EventSource or fetch readers.
  • Function Calling & Tool Streaming: Dynamically intercepting tool calls during active streaming to execute backend workflows.
  • Reverse Proxy Buffering Controls: Disabling Nginx buffer queues with X-Accel-Buffering to prevent chunk clumping.

1. Why Streaming? Anatomy of Latency in Generative Systems

As response lengths grow in large language model applications, total generation time can easily reach 5 to 15 seconds. Waiting for complete completion before dispatching a single JSON payload creates severe perceived friction.

Server-Sent Events (SSE) coupled with FastAPI async generators push individual token chunks the millisecond they are generated by Gemini, slashing Time to First Token (TTFT) to less than 250ms.

2. End-to-End Async Streaming with FastAPI & Google GenAI

A production-grade implementation streaming token chunks via FastAPI's `StreamingResponse`:

gemini_stream.py
import json
from typing import AsyncGenerator
from fastapi import FastAPI, HTTPException
from fastapi.responses import StreamingResponse
from google import genai
from google.genai import types

app = FastAPI(title="Gemini Streaming API")
ai_client = genai.Client()

async def generate_gemini_stream(prompt: str) -> AsyncGenerator[str, None]:
    try:
        response = ai_client.models.generate_content_stream(
            model="gemini-3.7-flash",
            contents=prompt,
            config=types.GenerateContentConfig(
                temperature=0.3,
                max_output_tokens=2048,
            )
        )
        for chunk in response:
            if chunk.text:
                payload = json.dumps({"text": chunk.text})
                yield f"data: {payload}\n\n"
    except Exception as exc:
        err_payload = json.dumps({"error": str(exc)})
        yield f"data: {err_payload}\n\n"

@app.get("/api/chat/stream")
async def chat_stream_endpoint(q: str):
    if not q.strip():
        raise HTTPException(status_code=400, detail="Query cannot be empty.")
    
    headers = {
        "Content-Type": "text/event-stream",
        "Cache-Control": "no-cache",
        "Connection": "keep-alive",
        "X-Accel-Buffering": "no",
    }
    return StreamingResponse(generate_gemini_stream(q), headers=headers)

3. Combining Streaming with Function Calling

When users trigger actions requiring live database queries or external APIs, the stream generator inspects `chunk.function_calls`. Upon receiving a tool invocation, the backend executes the corresponding Python function and streams the final synthesis back into the existing client connection.

4. Production Pitfalls: Nginx Buffering and Cloudflare Timeouts

When deployed behind Nginx or Cloudflare, reverse proxies may attempt to buffer packets, breaking the smooth token-by-token effect. Always attach `X-Accel-Buffering: no` and ensure reverse proxy timeouts (`proxy_read_timeout`) are set to at least 120 seconds.

Frequently Asked Questions

Should I use WebSockets or Server-Sent Events (SSE)?

For unidirectional text and AI token streaming, SSE is lighter, HTTP/2 multiplexing-native, and significantly easier to secure. WebSockets are reserved for bidirectional voice and live video streaming.

Does streaming increase API token costs?

No. Pricing is strictly calculated on input and output token volumes; streaming has zero surcharge.

Verified Documentation & Sources

Related Technical Guides

Deepen your understanding with these closely related production architectures and tutorials: