Guide the AI
Choose the spoken language or automatic detection, identify speakers, and add context or key terms for names and specialist vocabulary.
Transcribe audio to text or convert video to text with AI transcription. Review uncertain phrases against the source media, rename speakers, and export TXT, Markdown, SRT, or VTT.
No credit card required
TRANSCRIPTION
Upload one audio or video file and follow its progress from your browser.
HOW IT WORKS
Guide the AI before transcription, then check uncertain phrases against the source media, save your corrections, and export the latest revision.
Basic converter
File → AI-generated text → Download
Transcribe to Text
Guide → Transcribe → Review → Save → Export
Choose the spoken language or automatic detection, identify speakers, and add context or key terms for names and specialist vocabulary.
Upload a supported audio or video file. AI generates a structured transcript draft with timestamps and speaker labels when enabled.
Use the low-confidence review queue to play the matching moment, edit the segment, mark phrases reviewed, and rename speakers.
Corrections save as a new revision. Export the latest transcript as TXT, Markdown, SRT, or VTT.


LANGUAGE COVERAGE
Choose the recording’s spoken language, or use automatic detection when you are unsure.
supported languages
English includes General, Australia, United Kingdom, and United States options.
TRANSCRIPTION CAPABILITIES
A practical workspace for supported media, synchronized review, saved corrections, and reusable transcript exports.


TRANSCRIPTION USE CASES
A meeting transcript creates a written reference for decisions, project context, and follow-up after the conversation.
An interview transcript makes recorded answers easier to review, quote, compare, and organize during research or reporting.
A podcast transcript can support show notes, episode archives, accessibility work, and content repurposing.
A lecture transcript gives students and educators a text reference for classes, webinars, and training recordings.
A video transcript turns spoken dialogue into source text for scripts, editing notes, captions, and content review.
A customer call transcript helps sales, support, and research teams revisit questions, feedback, and conversation context.
1 credit equals 1 minute. Every credit you receive or purchase is yours until you use it.
Free credits for your first transcription, or predictable capacity with monthly and annual billing.
TRANSCRIPTION FAQ
Choose a supported audio file, set its spoken language and optional speaker labels, then start transcription. When processing finishes, open the review workspace to listen, correct segments, rename speakers, and export the latest saved revision.
Yes. Upload a supported MP4, MOV, AVI, MKV, MPEG/MPG, or WebM file and the service transcribes its spoken audio. You can review the result against the source video during the media retention window.
Supported files are AAC, AVI, FLAC, M4A, MKV, MOV, MP3, MP4, MPEG/MPG, OGG, OPUS, WAV, and WebM, up to 5 GB per file. A matching filename extension and media content type are required.
Audio clarity, microphone quality, background noise, accents, specialist terms, and overlapping speech can affect the result. The review queue groups lower-confidence phrases so you can replay and correct them; confidence is not presented as an accuracy guarantee.
Files upload directly to private, user-scoped storage. Source media remains available for review for up to seven days from task creation and can be deleted immediately. After deletion or expiry, the transcript, saved edits, history, and exports remain available without playback.
Transcribe to text means turning the spoken audio in a recording into written text. Upload a supported audio or video file, choose the spoken language and optional speaker settings, then review the saved transcript before exporting it.
Processing time depends on the recording length, media format, and audio complexity. The workbench shows the task status while transcription continues in the background; it does not promise a fixed completion time.
You can choose a known spoken language or use automatic detection when you are unsure. The language picker follows the languages currently supported by the transcription workflow, and the result should still be reviewed when the recording is difficult.
Yes. You can let the workflow detect speakers or provide a speaker count range, then rename speakers while reviewing the transcript. Clear turn-taking produces more useful labels than overlapping speech.
Yes. Add optional context or key terms before transcription when a recording contains specialist language. After processing, open the transcript, replay the source media, edit segments, correct low-confidence phrases, and save the latest revision before exporting.
Completed transcripts can be exported as TXT, Markdown, SRT, or VTT. TXT and Markdown are useful for documents and notes, while SRT and VTT preserve timing for subtitle workflows.
New accounts receive 10 one-time credits after registration. One credit represents one minute of transcription usage; paid plans and credit packs add more capacity, and purchased credits do not expire.
Upload a supported recording, review uncertain phrases against the source media, and export the saved transcript in the format you need.