Skip to content

Text to dialogueAvailable now

Scripts with more than one voice

Write a conversation as “Name: line”, give each speaker a voice and each line a delivery, and Vocla voices and mixes it into one file with the loudness evened out. Podcast intros, radio ads, sample calls and interviews, Arabic first.

  • Paste a whole script
  • Seven delivery styles per line
  • One mixed file

A real dialogue, line by line

نورةVoice: Noura
راشدVoice: Rashid
0:00 / 0:26

Real output, unedited · voices by Vocla
One Text to Dialogue generation: five Arabic lines, two library voices (Noura and Rashid), a style per line, one 0.6 s pause and quick-reply overlap on. Line times are the ones the API returned; the MP3 is re-encoded smaller for the web.

delivery styles, from natural to whisper
7
ready-made templates to start from
4
mixed file, loudness evened across voices
1
lines per script
200

How it works

From script to finished conversation

Dialogue is a mode of text to speech, so it uses the same voices, pronunciation dictionary and credits.

  1. Step 1

    Write or paste the script

    One line per turn as Name: line. Paste a whole script and every speaker is found for you.

  2. Step 2

    Give each speaker a voice

    Pick from the library, with your workspace's brand voices listed first.

  3. Step 3

    Set delivery and timing

    A style for each line, a pause after any line, a global gap and quick-reply overlap.

  4. Step 4

    Generate and listen

    Vocla voices every line, evens out the loudness and mixes one file. The player highlights each line as it plays.

Try it

Type Name: line and see the script split

This box follows the editor's rules and templates. It shows how Vocla reads a script; no audio is made here.

One line per turn. Add a style as [whisper] before the text or (calm) after the name. A line without a name continues the last speaker.

Start from a template

How Vocla reads it

5 lines · 2 speakers · 279 characters

  1. AnnouncerExcitedLooking for a coffee that actually wakes you up?
  2. LaylaHappyI tried Fajr Coffee this morning… honestly, I’m never going back!
  3. AnnouncerNaturalFajr Coffee — roasted fresh every day from the finest Yemeni beans.
  4. LaylaNaturalOrder in the app now and get twenty percent off your first order.
  5. AnnouncerCalmFajr Coffee. Your day starts here.

Dialogue is priced like text to speech: credits come from the characters of all lines and the model you choose, and are shown before you generate.

Delivery and timing

Direct every line

A conversation lives in its rhythm. These controls set how each line is said and when the next one starts.

  • A style for each line

    Natural, excited, calm, whisper, serious, happy or sad, set line by line.

  • A pause after any line

    Hold a beat after a line. A pause you set is kept exactly, up to 5 seconds.

  • A global gap

    The default silence between turns for the whole script, 0.35 seconds unless you change it (up to 3).

  • Quick-reply overlap

    A short reply from the other speaker can start up to 150 ms sooner, so quick back-and-forth sounds natural.

Why Vocla dialogue

Several voices, one finished file

No more generating lines one by one and lining them up in an editor.

  • Brand voices first

    When you cast speakers, your workspace's brand voices are offered first, so every production sounds like you.

  • Even loudness

    Each take is trimmed and levelled, then the mix is mastered, so no voice is louder than the other.

  • Follow along

    The player highlights each line in sync with the audio, and every line keeps its time in the mix.

  • Templates to start from

    Podcast intro, radio ad, customer call and interview scripts in Arabic and English, ready to edit.

Output

What you get

  • MP3The mixed conversation, on every plan
  • WAVLossless audio on Professional and Enterprise
  • LibrarySaved with the rest of your audio

For teams

One cast, the whole team

Dialogue uses your workspace's voices and rules, so everyone produces with the same brand voices and budgets.

  • RolesEditors write and generate; viewers can listen to what the team has made.
  • ApprovalsWhen your workspace requires approval, downloads wait for an approved review.
  • Credit capsMember credit caps apply to dialogue like any other generation.
  • Client tagsTag a dialogue with a client for clean usage reports.

Plans and credits

Priced like text to speech

Credits come from the number of characters in all lines and the model's rate, reserved when you generate and shown before you do. If a generation fails, its credits are refunded.

Priced like text to speech
Plans and creditsMonthly creditsWAV exportCommercial use
Free10,000Not includedNot included
Professional80,000IncludedIncluded
EnterpriseCustom volumeIncludedIncluded

Credits

How it's priced
Per character, like text to speech
The demo above
233 characters, 280 credits
Lines per script
Up to 200
Characters per line
Up to 2,000
Compare plans

Guide

How to make an AI dialogue in Arabic

What text to dialogue does

Text to dialogue turns a script with several speakers into one audio file. You write each turn as Name: line, cast a voice for each name, and Vocla voices every line in its speaker's voice, trims and levels the takes, and lays them on a timeline with the pauses you set.

Because it is part of text to speech, dialogue uses the same library voices, your workspace's brand voices, the pronunciation dictionary and number reading. Names you taught Vocla to say are said the same way in every line.

Writing a script that sounds natural

Write the way people speak: short sentences, one idea per line, and replies that answer the line before. Give a quick reaction its own line (“Really?”) rather than burying it in a long one; with quick-reply overlap on, short replies land a little sooner and the exchange feels alive.

Use styles sparingly. A whisper or an excited line stands out because the lines around it are natural. Add a pause after a question or before a punchline.

Choosing voices for a conversation

Pick voices that contrast: a female and a male voice, or two voices of different ages or dialects, so listeners always know who is speaking. For a brand, cast your brand voices first so every podcast intro and ad sounds consistent.

Podcast intros, radio ads and sample calls

The four templates are a quick start: a podcast intro with two hosts, a radio ad with an announcer and a customer, a customer-service call and an interview. Replace the names and lines with your own and keep the rhythm.

For training, sample calls with an agent and a customer in distinct voices are easy to update when a script changes, with no recording session.

Questions about Text to dialogue

What is text to dialogue?

A mode of text to speech that voices a script with several speakers and mixes it into one audio file, with a voice for each speaker and a delivery for each line.

How do I write the script?

One line per turn as Name: line. You can paste a whole script and every speaker is found. Add a style with [whisper] before the text or (calm) after the name.

Which delivery styles are there?

Natural, excited, calm, whisper, serious, happy and sad, chosen line by line.

Can I control the pauses?

Yes. Set a pause after any line (up to 5 seconds, kept exactly), a global gap between turns (0.35 seconds by default, up to 3) and quick-reply overlap for short answers.

Which voices can I use?

Any voice from the library, with your workspace's brand voices listed first, and your cloned voices on plans with voice cloning (your own voice, or one you have permission to use).

How much does a dialogue cost?

It is priced like text to speech, by the characters of all lines and the model you choose. The demo on this page, 233 characters, used 280 credits. The cost is shown before you generate.

Do the voices sound equally loud?

Yes. Each take is trimmed and levelled before mixing, then the whole file is mastered, so voices sit at the same loudness.

Can the speakers talk over each other?

Only slightly: with quick-reply overlap, a short reply can start up to 150 ms sooner. Each line is voiced separately, so there is no real crosstalk.

Is it only for Arabic?

No. Arabic comes first, but dialogue works with the library's English voices too, and the templates come in Arabic and English.

Are there templates?

Yes: podcast intro, radio ad, customer call and interview, each in Arabic and English.

What file do I get?

One mixed MP3, plus WAV on Professional and Enterprise. It is saved in your Library, and the player highlights each line as it plays.

Start in minutes

Give your content a voice today

Start free with 10,000 credits every month. No card required.

  • 10,000 free credits every month
  • No card required
  • Arabic and English interface