Skip to content

Speech to textAvailable now

Turn Arabic audio into text you can use

Upload a recording or a video. Vocla detects the language and who speaks when, writes every word with its time and can translate the transcript. Then you edit, search and export it in the format you need.

  • Speakers detected automatically
  • Word-level timestamps
  • Word, SRT, VTT, TXT, JSON and CSV

A real transcript and its translation

Speaker 1Speaker 2Detected: English
0:00 / 0:36

Transcript

Real output, unedited · voices by Vocla
Two-speaker clips voiced with Vocla library voices, uploaded and transcribed by Vocla with language and speakers on auto. Lines, speakers, word timings and the translation are exactly as returned.

export formats, from Word to CSV
6
credits per started minute of audio
40
per file on Professional
120 min
language and speaker detection
Auto

How it works

From recording to transcript in four steps

Most files are ready in a few minutes, and you see the cost before you start.

  1. Step 1

    Upload audio or video

    Drop in an interview, a lecture, a meeting or a video, up to 500 MB per file.

  2. Step 2

    Pick the language, or let Vocla detect it

    Arabic, English and French are recognised directly, other languages through AI detection. Add names and terms so they are spelled your way.

  3. Step 3

    Review in the editor

    Playback follows the text word by word. Fix a word, rename a speaker, split or merge lines, or find and replace across the file.

  4. Step 4

    Translate and export

    Translate the transcript side by side, then download Word, SRT, VTT, TXT, JSON or CSV with the original, the translation or both.

Accurate by design

Built for real Arabic recordings

Interviews, lectures and calls rarely come with one clear voice. Vocla keeps track of who said what and when.

  • Who spoke when

    Speakers are detected automatically and every line is labelled, so an interview reads like a script. Rename Speaker 1 to a real name once and it changes everywhere.

  • Every word has a time

    Word-level timestamps drive the editor's highlighting, precise subtitle timing and karaoke captions.

  • Arabic and its dialects

    Pick Modern Standard Arabic or a dialect, or leave it on auto: Arabic, English and French are detected directly, other languages through AI detection.

  • Names spelled your way

    Add a list of names and terms (people, places, brands) to improve how they are recognised. In the demo, the podcast's name and both hosts' names came out as listed.

The editor

Edit it like a document, check it like a recording

Everything happens against the audio, so checking a line takes a second.

  • Synced playback

    The line being spoken and the word inside it are highlighted as the audio plays.

  • Click to seek

    Click any line or word to hear it from there.

  • Inline editing

    Correct a word where it is. Timings stay attached to the line.

  • Speaker rename

    Turn Speaker 1 and Speaker 2 into real names for the whole transcript.

  • Split and merge

    Cut a long line where a new sentence starts, or join two short ones.

  • Find and replace

    Fix a spelling or a name across the whole file in one go.

Translation

Translate the transcript, side by side

Translate English to Arabic, Arabic to English and more. Each line keeps its time, so the translation exports to subtitles as easily as the original, or next to it in a bilingual file.

Translation adds 20 credits per minute.

From the demo: the Arabic transcript and its English translation
OriginalTranslation
  • أهلاً بكم في حلقة جديدة من بودكاست أصوات المدينة ضيفتنا اليوم مريم مصممة تطبيقات من جدة.

    Welcome to a new episode of the Aswat Al-Madina podcast. Our guest today is Maryam, an app designer from Jeddah.

  • أهلاً فهد وشكراً على الاستضافة.

    Hello Fahad, and thank you for hosting me.

  • حدثينا عن بدايتك كيف دخلت عالم التصميم؟

    Tell us about your beginnings. How did you enter the world of design?

  • بدأت قبل خمس سنوات بتطبيق صغير لمقهى في حينا ومن هناك كبرت الفكرة.

    I started five years ago with a small app for a coffee shop in our neighborhood, and from there the idea grew.

Exports

Take the transcript anywhere

Six formats, each with the original, the translation or both. The excerpts below are the real files exported from the demo.

Numbered cues with start and end times, wrapped for the screen, with Arabic marked right to left.

First lines of the exported file
1
00:00:00,400 --> 00:00:02,700
Speaker 1: Welcome back to the show today.

2
00:00:02,700 --> 00:00:05,500
Speaker 1: We're talking about
starting a small business.
  • Original, translation or both
  • Speaker names on or off
  • Timestamps on or off

Use cases

Who uses speech to text

Anyone who records people talking and needs the words in writing.

For teams

Transcripts that fit how your team works

Speech to text runs inside your Vocla workspace, with the same roles, approvals and budgets as the rest of the product.

  • RolesViewers can open, search and export transcripts; editors upload, edit, translate and render.
  • ApprovalsWhen your workspace requires approval, caption renders wait for an approved review before anyone downloads them.
  • Credit capsMember credit caps cover transcription, translation and caption renders.
  • Client tagsTag a transcript with a client and its charges and caption renders follow, for clean usage reports.

Plans and credits

Clear prices, clear limits

Transcription is charged per started minute, from the file's length, and the cost is shown before you start. If a transcript fails, its credits are refunded.

Clear prices, clear limits
Plans and creditsLongest fileTranslationAll six exports
Free10 minutesIncludedIncluded
Professional120 minutesIncludedIncluded
Enterprise300 minutesIncludedIncluded

Credits

Transcription
40 credits per started minute
Translation
+20 credits per minute
Upload size
Up to 500 MB per file
Example: a 30-minute interview
1,200 credits (1,800 with translation)
Compare plans

Guide

How to transcribe Arabic audio with AI

How Arabic speech to text works

Speech to text listens to a recording and writes down what is said, with the time of every word. Vocla first finds the language (or uses the one you choose), then recognises the speech, splits it into lines at pauses and sentence ends, and works out which speaker says each line.

The result is not a wall of text: it is a transcript you can play, search and edit, with each line tied to its moment in the audio. That is what makes Arabic transcription useful for subtitles, quotes and research, not only for reading.

Getting the most accurate Arabic transcription

Clear audio matters more than anything else. Record close to the speaker, avoid music under speech, and give each person their own turn. If you know the dialect, choose it; if the clip is short or mixes languages, choosing the language yourself is more reliable than auto-detect.

Names are where any recogniser struggles, so add the people, places and brands in the recording to the names list before you start. After that, a pass in the editor with find and replace usually takes a few minutes, even on long files.

Transcripts, subtitles or captions?

A transcript is the full text with speakers and times, for reading, quoting and searching. Subtitles are the same words cut into short timed cues (SRT or VTT) for a player. Captions are subtitles burned into the picture, so they show everywhere, including on social feeds with the sound off.

In Vocla they all come from one transcript: fix a word once and the Word file, the subtitles and the captioned video all use the corrected line.

Translating a transcript

Translation runs line by line with the surrounding lines as context, and keeps each line's time. You can read it next to the original, export it alone, or export both in one bilingual file, which is handy for reviewers who read only one of the two languages.

Translate English interviews into Arabic for an Arabic audience, or Arabic recordings into English for partners and research. Translation adds 20 credits per minute to the 40 of transcription.

What a transcript costs

Transcription uses 40 credits per started minute, taken from the file's length before it starts and refunded if it fails. A 30-minute interview uses 1,200 credits, or 1,800 with a translation. The Free plan's 10,000 monthly credits cover about four hours of transcription, with files up to 10 minutes each.

Questions about Speech to text

How accurate is Arabic speech to text?

On clear recordings the words are usually right, with Arabic spelling (hamzas, ة and ى) restored; the demo on this page is published unedited so you can judge for yourself. Expect to check names and add some punctuation, since transcripts come without tashkeel and punctuation inside a line can be sparse. The editor makes that a quick pass.

Which languages and dialects can I transcribe?

Arabic, including dialects you can choose, plus English, French and other languages. On auto, Arabic, English and French are detected directly and other languages through AI detection.

Does it detect different speakers?

Yes. Speakers are detected automatically and every line is labelled Speaker 1, Speaker 2 and so on. You can rename speakers and move lines between them in the editor.

Can I transcribe a video, not just audio?

Yes. Upload audio or video; Vocla transcribes the soundtrack. The same transcript can then become burned-in captions with AI captions.

How long can a file be?

Up to 10 minutes per file on Free, 120 minutes on Professional and 300 minutes on Enterprise, and up to 500 MB per upload.

How much does transcription cost?

40 credits per started minute of audio, plus 20 credits per minute if you translate. The cost is shown before you start and refunded if the transcript fails.

Can I translate the transcript?

Yes, for example English to Arabic or Arabic to English. The translation keeps each line's time and appears side by side with the original, and you can export the original, the translation or both.

Which export formats are there?

Word (DOCX) with proper right-to-left Arabic, SRT and VTT subtitles, plain text, JSON with word timings, and CSV that opens correctly in Excel.

Do Arabic transcripts open correctly in Word and Excel?

Yes. The Word file uses right-to-left paragraphs for Arabic lines, and the CSV is saved as UTF-8 with a byte-order mark so Excel shows Arabic properly.

How do I improve how names are recognised?

Add them to the names and terms list when you start the transcript. Vocla uses the list to recognise them and writes a matching word the list's way.

Is there live, real-time transcription?

No. Speech to text works on uploaded files, usually ready within a few minutes.

Can my team work on transcripts together?

Yes. Transcripts live in your workspace: viewers can read and export, editors can edit and translate, and member credit caps and client tags apply.

Start in minutes

Give your content a voice today

Start free with 10,000 credits every month. No card required.

  • 10,000 free credits every month
  • No card required
  • Arabic and English interface