Skip to content

About

Azure Voice Live demo: Browser audio via FastAPI + Terraform

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

Speech-to-Speech (Azure Voice Live) Demo

This is a portfolio demo showing end-to-end skills across Azure, Terraform, Python (FastAPI), and AI realtime speech.

Flow: Browser microphone → 24 kHz mono audio → WebSocket → FastAPI relay → Azure Voice Live → audio deltas → browser playback.

What this demonstrates

  • Azure + Terraform: provisions Azure AI Services (Foundry/AIServices), Key Vault, a Linux Web App with system-assigned managed identity + Key Vault secret reference, and a Log Analytics workspace with diagnostic settings.
  • Python: FastAPI app serving a small web UI and a WebSocket relay.
  • AI: realtime Voice Live session updates, audio streaming, and response audio playback.

Architecture

Azure resources provisioned by Terraform in infra/:

flowchart LR
    User["User browser"]

    subgraph Azure["Azure resource group"]
        Plan["App Service plan"]
        App["Linux Web App<br/>FastAPI + static UI"]
        KV["Key Vault<br/>secret: foundry-api-key"]
        Foundry["Azure AI Services<br/>Foundry / Voice Live"]
        LAW["Log Analytics workspace"]
    end

    User -->|https + wss| App
    App -->|Key Vault reference| KV
    App -->|wss + api-key header| Foundry
    App -.->|diagnostic setting| LAW
    Foundry -.->|diagnostic setting| LAW
    Plan --- App
Loading

Realtime audio flow for a single session:

sequenceDiagram
    autonumber
    participant UI as Browser
    participant API as FastAPI relay
    participant VL as Azure Voice Live

    UI->>UI: Capture mic, resample to 24 kHz, convert to PCM16
    UI->>API: WebSocket open with Origin header
    Note over API: Validate Origin against allowlist<br/>or PUBLIC_WEB_URL
    API->>VL: WS connect with api-key request header
    Note over API,VL: Secret travels in the header, never in the URL
    API->>VL: session.update - modalities, voice, server VAD
    UI->>API: binary PCM16 frames
    API->>VL: input_audio_buffer.append - base64 PCM16
    VL-->>API: response.audio.delta - base64 PCM16
    API-->>UI: forwarded JSON deltas
    UI->>UI: Decode and schedule on AudioContext for playback
Loading

Prereqs

  • Python 3.12+
  • Azure CLI logged in (az login)
  • Terraform 1.14+

Provision Azure infrastructure (Terraform)

From infra/:

cd infra
copy terraform.tfvars.example terraform.tfvars
# edit terraform.tfvars and set your subscription_id
terraform init
terraform apply

Run locally (after provisioning)

  1. Create a local env file:
cd ..
copy .env.example .env
  1. Configure .env (repo root)

The API loads environment variables from .env at the repository root (see src/api/config.py).

Set these variables:

  • VOICE_LIVE_WS_ENDPOINT: the base Voice Live WebSocket URL (must include api-version=...).
    • After Terraform provisions resources, copy the exact value:
cd infra && terraform output -raw voice_live_ws_endpoint
  • Paste it into .env as a single line, for example:
VOICE_LIVE_WS_ENDPOINT=wss://<your-account-subdomain>.services.ai.azure.com/voice-live/realtime?api-version=2025-10-01
  • FOUNDRY_API_KEY: your Azure AI Services / Foundry key.
    • Recommended: print it from Key Vault using the helper script (requires az login + Terraform outputs):
.\scripts\get-foundry-key.ps1
  • Copy the FOUNDRY_API_KEY=... line into .env.

  • VOICE_LIVE_MODEL (optional): overrides the default model in src/api/config.py.

  • VOICE_LIVE_INSTRUCTIONS (optional): full Voice Live system prompt (session.update). If set in .env or App Service, it replaces the default string in src/api/config.py (otherwise edit that default in code).

  • ENVIRONMENT (optional): defaults to local. Used for small behavioral differences between local and deployed environments.

  • HTTP_RATE_LIMIT (optional): defaults to 60/minute (SlowAPI) for basic HTTP abuse protection on / and /config.

  1. Install + run:
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -e ".[dev]"
python -m uvicorn api.main:app --host 127.0.0.1 --port 8000

Then open http://127.0.0.1:8000/ and click Start.

Deploy to Azure (run remotely)

After terraform apply, deploy the app to the provisioned Linux Web App:

.\scripts\deploy-azure-webapp.ps1

When it finishes, open the web_app_url output:

cd infra && terraform output web_app_url

Destroy Azure resources (minimize cost)

Everything under this stack (resource group, App Service plan, Web App, Cognitive Services account, Key Vault, etc.) keeps incurring charges until you delete it. When you are done experimenting, tear it down from infra/:

cd infra && terraform destroy

About

Azure Voice Live demo: Browser audio via FastAPI + Terraform

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages