Back to Skills

text-to-speech

Converts text to speech audio using OpenAI TTS API. Use when users request audio versions of text or want responses read aloud.

3starsUpdated 3/25/2026

Security Assessment

Safe(100/100)
Security Score100/100

About text-to-speech

The 'text-to-speech' skill is designed to convert written text into audio, making it easier for users to consume content in a hands-free manner. This is particularly useful in scenarios where users require audio versions of text, such as when they need to listen to content on the go, or when responding to requests for audio-based content. By leveraging OpenAI’s TTS (Text-to-Speech) API, this skill provides natural and intelligible speech synthesis for a wide variety of texts, allowing for seamless integration in different use cases.

Main features of this skill include the ability to generate speech in multiple voices, adjust the speech speed, and select the quality of the audio model. Users can opt for a regular audio file output or a voice message format, which is tailored specifically for use within messaging platforms like Telegram. The voice options range from neutral tones to expressive, deeper, or softer voices, and the skill also supports various audio formats like MP3, OGG, AAC, FLAC, and WAV. The command line interface is simple to use, allowing customization through flags like --voice, --speed, and --model, offering flexibility depending on the user’s preferences.

This skill is ideal for individuals or systems that require voice interaction, such as customer support bots, accessibility tools, and content creators who want to provide audio alternatives for their written materials. Additionally, it's a great tool for users who prefer voice communication or require auditory content for multitasking or accessibility reasons. Whether it’s answering questions, providing summaries, or reading articles aloud, this skill enhances user experience by offering text-to-speech functionality in various formats.

FAQ

How do I use the text-to-speech skill?

You can use the skill by running the command 'telclaude tts' followed by the text you want to convert to speech. Use the --voice-message flag for voice message replies in Telegram, or omit it for regular audio files.

Can I adjust the voice used in the speech?

Yes, the skill allows you to choose from multiple voice options, including alloy, echo, fable, onyx, nova, and shimmer, each with different tonal qualities. The default voice is 'alloy'.

What are the available audio formats?

You can choose from several audio formats such as mp3, opus, aac, flac, and wav. The --voice-message flag outputs the audio in OGG format (Opus), suitable for Telegram.

Can the speech speed be customized?

Yes, the speech speed can be adjusted from 0.25 to 4.0, with the default speed set to 1.0.

Are there any limitations with the audio quality?

There are two quality models: 'tts-1' for standard quality and 'tts-1-hd' for higher quality. The 'tts-1-hd' model produces higher-quality audio but is slightly slower than the standard model.

Install text-to-speech

Download and extract the skill files to your .claude/skills/ directory.

Quick Setup:

  1. Copy the skill folder to .claude/skills/
  2. Claude will automatically detect and use the skill