Use this skill when building real-time, bidirectional streaming applications with the Gemini Live API. Covers WebSocket-based audio/video/text streaming, voice activity detection (VAD), native audio features, function calling, session management, ephemeral tokens for client-side auth, and all Live API configuration options. SDKs covered - google-genai (Python), @google/genai (JavaScript/TypeScript).
Detected risks:
The gemini-live-api-dev skill enables developers to build **real-time, interactive applications** using the Gemini Live API over WebSockets. It addresses the challenge of creating low-latency, bidirectional audio, video, and text communication, allowing applications to process continuous streams for immediate responses. By leveraging this skill, developers can integrate human-like conversational capabilities into their applications, ensuring seamless real-time interactions between users and AI. The skill also covers secure authentication using ephemeral tokens and session management for context persistence and efficient resource handling.
Key capabilities include bidirectional audio streaming for real-time conversations, video streaming with synchronized camera or screen frames, text input/output within live sessions, and automatic audio transcription. Voice Activity Detection (VAD) ensures natural interruption handling, while native audio features like configurable thinking levels allow the AI to simulate thought processing. Function calling and Google Search grounding extend functionality to synchronous tool usage and contextually accurate responses. Session management features such as context compression, session resumption, and GoAway signals help maintain continuity and reliability during interactions.
This skill is particularly useful for developers building **real-time AI chatbots, collaborative audio/video platforms, and interactive customer support applications**. Target users include Python and JavaScript/TypeScript developers utilizing the google-genai and @google/genai SDKs. It is also valuable for teams seeking to integrate AI capabilities into third-party platforms via partner integrations like LiveKit, Pipecat, Fishjam, Vision Agents, Voximplant, and Firebase AI SDK. The gemini-live-api-dev skill provides a comprehensive foundation for creating responsive, immersive, and secure real-time AI experiences.
The skill supports Python through the google-genai SDK and JavaScript/TypeScript via the @google/genai SDK. Legacy SDKs are deprecated.
No, the Live API currently only supports WebSockets. WebRTC support is available through partner integrations such as LiveKit or Pipecat.
Input audio must be raw PCM, little-endian, 16-bit, mono at 16kHz. Output audio is raw PCM, little-endian, 16-bit, mono at 24kHz.
Yes, the skill supports sending camera or screen frames alongside audio for real-time video streaming.
Yes, models like gemini-2.5-flash-native-audio-preview-12-2025 and gemini-live-2.5-flash-preview are deprecated. Use gemini-3.1-flash-live-preview for all current Live API applications.
Quick Setup:
.claude/skills/Repository
google-gemini/gemini-skills