Skip to content
Guide

Add a Flutter voice agent over WebSocket with Wixzel Voice

At the end your Flutter app lets a user talk to a voice agent you built, over a WebSocket, with no SIP trunk and no phone number. Your server mints a single-use session, so your API key never reaches the app, and the wixzel_voice Dart SDK speaks the protocol. Expect about an hour, most of it wiring the microphone and speaker. A session is billed like a phone call on the same engine: $0.0466 per connected minute on the Gemini Live engine, about $0.19 for a four-minute session, per second from a prepaid balance, with no telephony cost.

By · Published · Updated

Takes
About an hour
Engine time
$0.0466 per connected minute

What do you need before you start?

  • A Wixzel Voice API key with credit, kept on your server.
  • A server you control to mint sessions. The example uses Node.js and Express; any language that can make an HTTPS request works.
  • A Flutter app with a recording plugin that streams 16-bit PCM, ideally at 8 kHz mono with echo cancellation, and a player that accepts raw PCM.

How do you build this on Wixzel Voice, step by step?

  1. 01Create the agent once

    This agent uses Gemini Live, one realtime model that listens, reasons and speaks:

    shell
    curl https://api.voice.wixzel.com/v1/agents \
      -H "Authorization: Bearer $WIXZEL_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "name": "In-app assistant",
        "system_prompt": "You are the voice assistant inside the Acme app. Answer in one or two short sentences.",
        "opening_message": "Hi, what can I help you with?",
        "voice": {
          "realtime": { "model": "google/gemini-live", "voice": "Kore" }
        }
      }'

    Nothing about this agent is specific to the app: the same agent could answer a phone number. Keep its id on your server as AGENT_ID.

  2. 02Mint a session on your server

    The app never holds your API key. It asks your server, and your server asks the API for a session with POST /v1/realtime/sessions, here through the TypeScript SDK:

    server.mjs
    import express from 'express';
    import { WixzelVoice } from 'wixzel-voice';
    
    const wixzel = new WixzelVoice({ apiKey: process.env.WIXZEL_API_KEY });
    const app = express();
    
    // Replace with your own authentication. It must set req.userId.
    function requireUser(req, res, next) {
      req.userId = 'demo-user';
      next();
    }
    
    app.post('/voice/session', requireUser, async (req, res) => {
      const session = await wixzel.realtime.createSession({
        agent_id: process.env.AGENT_ID,
        max_duration_seconds: 300,
        metadata: { user_id: req.userId },
      });
      res.json(session);
    });
    
    app.listen(3000);

    The response carries the socket url, a client_secret that works once, for one minute, for this agent, and audio_format: "mulaw_8000". Minting costs nothing; billing starts when the app connects. max_duration_seconds caps the conversation, anywhere from 10 to 3600 seconds, 600 by default. metadata is echoed on the call record and never sent to the model. allowed_origins matters only for browsers, so a native app leaves it out.

  3. 03Connect from the app

    Add the SDK with flutter pub add wixzel_voice http, then open the session your server returned:

    voice_call.dart
    import 'dart:async';
    import 'dart:convert';
    import 'dart:typed_data';
    
    import 'package:http/http.dart' as http;
    import 'package:wixzel_voice/wixzel_voice.dart';
    
    /// Whatever your playback plugin offers, behind a small interface.
    abstract class PcmPlayer {
      void add(Int16List samples);
      void clear();
    }
    
    class VoiceCall {
      VoiceCall({required this.player, required this.onLine});
    
      final PcmPlayer player;
      final void Function(String role, String text) onLine;
    
      RealtimeConnection? _conn;
      StreamSubscription<Uint8List>? _mic;
      final List<int> _pending = [];
    
      /// [micBytes] is 16-bit little-endian PCM, 8 kHz mono, from your recording plugin.
      Future<void> start(Stream<Uint8List> micBytes) async {
        final res = await http.post(Uri.parse('https://api.example.com/voice/session'));
        final status = res.statusCode;
        if (status != 200) throw Exception('Could not start a voice session: HTTP $status');
        final session = RealtimeSession.fromJson(jsonDecode(res.body) as Map<String, dynamic>);
    
        final conn = RealtimeConnection.forSession(session);
        _conn = conn;
        conn.events.listen((event) {
          switch (event) {
            case RealtimeAudio(:final mulaw):
              player.add(mulawToPcm16(mulaw));
            case RealtimeAudioClear():
              player.clear(); // the user interrupted: drop queued audio
            case RealtimeTranscript(:final role, :final text):
              onLine(role, text);
            case RealtimeError(:final code, :final message):
              onLine('error', '$code: $message');
            case RealtimeEnded(:final reason):
              onLine('system', 'ended: $reason');
              _mic?.cancel();
            default:
          }
        });
        await conn.ready;
        _mic = micBytes.listen(_send);
      }
    
      void _send(Uint8List bytes) {
        final data = ByteData.sublistView(bytes);
        for (var i = 0; i + 1 < bytes.length; i += 2) {
          _pending.add(data.getInt16(i, Endian.little));
        }
        // 160 samples is 20 ms at 8 kHz, the frame size the server expects.
        while (_pending.length >= 160) {
          final frame = Int16List.fromList(_pending.sublist(0, 160));
          _pending.removeRange(0, 160);
          _conn?.sendAudio(pcm16ToMulaw(frame));
        }
      }
    
      Future<void> end() async {
        await _mic?.cancel();
        await _conn?.end();
      }
    }

    RealtimeConnection speaks the protocol and nothing else: it does not open the microphone or the speaker, which keeps the SDK free of plugin dependencies and leaves you to use the audio plugins you already have. pcm16ToMulaw and mulawToPcm16 convert between 16-bit PCM and what goes over the wire.

    Ask for microphone permission first: RECORD_AUDIO on Android and NSMicrophoneUsageDescription on iOS.

  4. 04Turn on echo cancellation, and expect phone-quality audio

    Turn on your platform's echo cancellation in the recording plugin. Without it the agent hears itself through the speaker and interrupts itself.

    Audio is G.711 µ-law at 8 kHz in both directions, the same as a phone line, because sessions run through the same media path as calls. That is fine for speech and narrower than wideband audio. Record at 8 kHz if your plugin allows it, or resample before pcm16ToMulaw, and play the agent's audio at 8 kHz mono.

  5. 05Read the session afterwards

    Every session is a call record with channel: "web":

    shell
    curl "https://api.voice.wixzel.com/v1/calls?channel=web&limit=5" \
      -H "Authorization: Bearer $WIXZEL_API_KEY"

    A web session is inbound, has no from or to, and carries its duration, cost_micros and your metadata; GET /v1/calls/{id} adds the transcript. It ends with the same callCompleted webhook a phone call sends. Web sessions are not recorded.

What does this build cost on Wixzel Voice?

Engine time for 4-minute sessions on the gemini-live engine at $0.0466 per connected minute: about $0.19 per session. There is no telephony cost, because there is no phone line.

Monthly engine cost for this guide’s build
Sessions per monthConnected minutesEngine cost per month
5002,000$93.20
2,50010,000$466
10,00040,000$1,864

A session costs what a phone call on the same engine costs, per second, from the same prepaid balance. Minting a session costs nothing, so a user who opens the screen and then denies the microphone costs nothing.

Prepaid, with no subscription and no free tier: add credit from $5 and usage spends it. Estimate your own mix.

Which Wixzel Voice limits apply to this build?

  • Phone-quality audio: G.711 µ-law at 8 kHz in both directions, not wideband.
  • The Dart SDK does not capture or play audio. Your app supplies the microphone and the speaker through plugins of its choice.
  • Sessions count toward the account's concurrent-call limit, five by default, alongside phone calls. The limit is raised on request.
  • A session lasts 10 minutes by default and at most an hour, and two minutes of silence ends it.
  • Web sessions are not recorded, and there is no second line to hand the user to a person.

Frequently asked questions

Is there a voice agent API for Flutter?
Wixzel Voice has one. Its Dart SDK, wixzel_voice on pub.dev, includes RealtimeConnection, which connects a Flutter or Dart app to a voice agent over a WebSocket. Your server mints the session with your API key; the app only ever holds a one-minute, single-use secret.
Do I need a phone number or a SIP trunk for an in-app voice agent?
No. Sessions in web and mobile apps run over a WebSocket with nothing to connect. A trunk and a number are only needed for calls to and from real phones.
Why does the agent keep interrupting itself?
It is hearing its own voice through the speaker. Turn on echo cancellation in your recording plugin, or test with headphones to confirm.
Can I use a different engine in the app?
Yes. The classic and sarvam pipelines run over the socket too: choose classic for a specific ElevenLabs voice, or sarvam for Indian languages. The price is the engine's own per-minute price, the same as on a phone call; Gemini Live, at $0.0466, is the one used here.
What happens when the account runs out of credit mid-session?
The app receives a session.warning with code low_balance and roughly how many seconds are left, then the session ends with reason insufficient_credits. A session that cannot be funded at all is refused with close code 4402.

Where to read more

  • Realtime: web and mobiledocs

    Minting sessions, the protocol, refusal codes and billing.

  • SDKsdocs

    The Dart and TypeScript SDKs, including the browser client.

  • Voice enginesdocs

    Realtime and composed voices, and the voice rosters.

  • Webhooksdocs

    callCompleted for web sessions, with your metadata.

APIs for agentic telephony

About Wixzel Voice

Wixzel Voice is a voice AI API for building AI agents that place and answer real phone calls over your own SIP trunk, or talk to people in your web and mobile apps. One API key and one prepaid balance cover every voice engine, billed per second of actual usage.