Speech to Text

Turn Khmer Audio into Accurate Text

Upload a recording or record in the browser and get back a timestamped transcript you can edit, label by speaker, and export as subtitles.

interview.mp3
00:12
00:00
Speaker 1

Welcome to Kiri TTS. Let me walk you through the recording.

00:04
Speaker 2

Sounds good — I have already uploaded the audio file.

00:09
Speaker 3

Great, the transcript is ready to export.

Export asSRTVTTJSONTXTCSV

Everything you need to transcribe audio

Speaker labels, accurate timings, an editor that keeps up, and exports that drop straight into your workflow.

Automatic speaker detection

Transcripts are split by speaker as they come back. Rename each one and give them a colour — the labels carry through to your exports.

Timestamped, editable lines

Every line keeps its own start and end time, aligned to the audio, and stays fully editable in the browser.

Subtitle-ready exports

Download as SRT, VTT, JSON, TXT or CSV. Cap the words per line so long sentences split into readable caption cues.

Trim before transcribing

Select just the section you need before you spend credits. Handles up to one hour of audio per run.

Projects and history

Group transcriptions into projects and come back to any of them later. Nothing is lost after you close the tab.

OpenAI-compatible API

Call the transcription endpoint from your own code. Existing OpenAI SDK calls work by changing the base URL.

How it works

From recording to subtitle file in three steps.

1

Upload or record

Drop in an audio file or record straight from your browser. Trim to just the part you need, up to one hour per run.

2

Get a timestamped transcript

Speakers are detected automatically and every line carries its own start and end time, aligned to the audio.

3

Edit and export

Fix any wording, rename speakers, then export as SRT, VTT, JSON, TXT or CSV.

Frequently Asked Questions

Yes. Kiri is built for Khmer first, and also transcribes English and other languages. Khmer audio returns Khmer script with correct word boundaries, not a romanised approximation.

Ready to transcribe your first recording?

Transcription uses the same credits as speech generation, so every plan includes it — the free one too.