Request a tool
All toolsGuidesMCP serverRequest a toolPlatformsCategories
AI Text-to-Speech Voiceover icon

AI Text-to-Speech Voiceover

Turn any script into a natural AI voiceover audio file (MP3/WAV/Opus/AAC). Pick a voice, speed, and format. Good for faceless videos, IVR, and narration.

5 from 3 reviews on Apify 31 runs on Apify $0.04 per voiceover ($40 / 1,000)
Run this in the cloudRun on Apify →

YouTube & Creator Tools

How it works

  1. 1
    Open it on Apify

    Hit Run on Apify — it opens the tool in the cloud, no install.

  2. 2
    Set the inputs

    Adjust text, texts, voice (sensible defaults are pre-filled).

  3. 3
    Click Run

    The tool runs on Apify’s cloud and collects the data for you.

  4. 4
    Export the results

    Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.

Pricing

$0.04 per voiceover = $40 per 1,000

You are charged forWhenPrice
Voiceover generatedOne generated voiceover (up to 5k chars).$0.04
Actor StartCharged when the Actor starts running. Number of events charged depends on Actor memory (one event per GB, minimum one event).$0.0000125

Pay-per-event pricing: you are billed per result, not per subscription — a run that returns nothing costs nothing beyond the start fee. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-08-08, and they are what you are actually charged.

Inputs

FieldWhat it doesType
textThe text to convert to speech. Long scripts are chunked and stitched automatically.string
textsArray of strings OR objects (uses script/scriptText/text/narration). One audio file per item.array
voiceAI voice.string
modeltts-1 (fast) or tts-1-hd (higher quality).string
formatOutput audio format.string
speedPlayback speed 0.25–4.0 (1.0 = normal).string
openaiApiKeyYour OpenAI key (TTS). Kept private.string
baseUrlOpenAI-compatible base URL. Default https://api.openai.com/v1.string

What you get

A structured dataset — each result includes fields like:

_demo_noticeaudioKeyaudioUrlcharacterschunksdurationSecondsformatindexmodeltextPreviewvoice

Export every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.

3 ready-to-run use cases

Faceless YouTube AI Voiceover from a Script

Paste a narration script and get a steady AI voiceover track for faceless YouTube videos, ready to drop under your B-roll. MP3 or WAV, adjustable pace.

IVR Phone Menu Voice Prompts - Text to Speech

Turn "press 1 for sales" menu text into clear IVR voice prompts for your phone system, exported as AAC. Telephony-ready greetings without studio time.

Batch Text to Speech: One Audio File per Line

Got a list of lines? Each one comes back as its own AI voiceover file, ideal for app prompts, UI sounds, and clip sets. Pick the voice, speed, and format.

Related tools in YouTube & Creator Tools

Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.

Video & Audio Transcriber — Word-Level + SRT/VTT iconYouTube & Creator Tools

Video & Audio Transcriber — Word-Level + SRT/VTT

Transcribe any video or audio URL into text with word-level timestamps and ready SRT, VTT, and TXT files. Auto language detection, batch mode.

3 use cases

AI Character Reference Bank iconYouTube & Creator Tools

AI Character Reference Bank

Turn a text description into a consistent character reference set (portrait, 3/4, full-body, expressions) plus a reusable character-bible prompt. No code.

2 use cases

YouTube Channel Scraper iconYouTube & Creator Tools

YouTube Channel Scraper

Scrape every video from any YouTube channel, no login or API key: views, dates, duration and thumbnails.

Ready to run — no setup

YouTube Comments Scraper iconYouTube & Creator Tools

YouTube Comments Scraper

Scrape YouTube comments and replies with no login or API key: text, likes, reply counts, authors and timestamps.

Ready to run — no setup

YouTube Playlist Scraper iconYouTube & Creator Tools

YouTube Playlist Scraper

Export every video in a YouTube playlist: title, channel, duration, views, publish date and URL. No API key, OAuth or quota.

Ready to run — no setup

YouTube Transcript Scraper iconYouTube & Creator Tools

YouTube Transcript Scraper

Extract YouTube transcript text, timestamped segments, SRT, VTT, languages and video metadata without an API key.

Ready to run — no setup

See all YouTube & Creator Tools →

AI Text-to-Speech Voiceover

Turns a block of text or a full script into a natural-sounding AI voiceover file. Pick a voice, set the speed, and choose MP3, WAV, Opus, or AAC. It's meant for the usual narration jobs: faceless videos, audiobooks, IVR prompts, explainer voiceovers.

How it works

The actor sends your text to an OpenAI-compatible TTS endpoint. Long scripts get split at sentence boundaries into chunks under ~3,500 characters, each chunk is synthesized separately, and the parts are stitched back into one file with ffmpeg (using stream copy, so there's no re-encode and no quality loss). Each finished audio file is saved to the run's key-value store and a row is pushed to the dataset.

Input

Nothing is strictly required by the schema, but in practice you need an openaiApiKey and at least one of text or texts. If neither is provided the run errors out.

FieldRequiredNotes
textone of text/textsThe script to voice, as a single string.
textsone of text/textsBatch mode. Array of strings, or objects keyed by script / scriptText / text / narration. One audio file per item.
voicenoalloy, echo, fable, onyx, nova, shimmer. Default onyx (deep male). nova and shimmer are female.
modelnotts-1 (fast, default) or tts-1-hd (higher quality, costs more on the OpenAI side).
formatnomp3 (default), wav, opus, or aac.
speednoPlayback speed from 0.25 to 4.0. Default 1.0. Values outside that range are clamped.
openaiApiKeyyes in practiceYour OpenAI key, used for the TTS call. Stored as a secret. Falls back to the OPENAI_API_KEY env var if set.
baseUrlnoAdvanced. Point at any OpenAI-compatible /audio/speech endpoint. Defaults to https://api.openai.com/v1.

Output

Each input item produces one audio file in the key-value store and one dataset record. The record includes audioKey and audioUrl (where to fetch the file), durationSeconds, characters, chunks (how many pieces the script was split into), plus the voice, model, and resolved format. Failed items get a record with ok: false and the error message instead of stopping the whole run.

Example

{
  "text": "Welcome back to the channel. Today we're looking at one of the strangest mysteries of the deep ocean.",
  "voice": "onyx",
  "model": "tts-1",
  "format": "mp3",
  "speed": 1.0,
  "openaiApiKey": "sk-..."
}

Pricing

$0.04 per voiceover — $40.00 per 1,000 — plus a $0.0000125 run-start fee each time a run begins. Flat rate: no volume tiers, no plan gates, no subscription. The OpenAI TTS usage is billed separately on your own key, so the only Apify-side cost is the per-voiceover fee above.

Batch mode is where that rate matters. Because texts accepts an array, one run can voice a whole week of scripts in a single start fee — 50 narration blocks cost $2.00 plus a fraction of a cent to start the run.

Notes

This actor calls OpenAI for synthesis, so it needs your own OpenAI API key. Individual chunks are capped at 4,000 characters before they're sent, which keeps each request within the model's per-call limit; there's no hard limit on total script length since long inputs are chunked and concatenated.