Every claim on this site — diarization, SRT, VTT, PDF, summaries, action items, translated subtitles that keep their timecodes — is easy to state and hard to verify from a marketing page. So instead of describing what convert.express produces, this page shows one real job, unedited, from upload to every artifact it generated. Nothing here is a mockup: it is a 4-minute audio clip run through convert.express once, with every output pasted in as-is.

The recording

A 4-minute excerpt from a real conversation, uploaded as WhatsApp Audio - 4min demo.mp4. Two people are talking; the topic is performing arts and social media. At upload, the settings were: Google Gemini 3.5 Transcribe (chosen for its in-model speaker diarization), speaker diarization on, AI summary on, AI action items on. Language was auto-detected.

File:      WhatsApp Audio - 4min demo.mp4
Model:     Gemini 3.5 Transcribe
Duration:  241 seconds (4:01)
Language:  auto-detected (English)
Options:   speakerDiarization=true, generateSummary=true, generateActionList=true

The diarized transcript

This is the core deliverable: not just text, but who said it. Gemini's in-model diarization split the recording into two speakers and labeled every turn — this is the raw diarizedSegments output, unedited.

Speaker A: I think that a really big problem and this is the voice because
the voice is a very personal thing that you expose to the audience.

Speaker B: Part of your body basically.

Speaker A: Yes, you know part of your body or of your mind. So it's very
personal so but if you choose to do it in your life, if you decide to
to start in to be a singer

Speaker B: Right.

Speaker A: you you you should you should try to expose yourself also in
the social media because it's the same. So I think it could could be a
a beautiful tirocinio and this is pre-work. I don't know what to say.

Speaker B: Like beautiful try, beautiful experimental

Speaker A: Training, yeah training, beautiful training to work.

Speaker B: In a real world.

Speaker A: So you start with yeah you start with the yeah? You improve
your yourself and your communication and and then you could have also
when when you are in social media you could also delete the post. If
you are if you if you if you are in in front of a real audience you
can't do it.

Speaker B: They impression there forever.

Speaker A: Yeah, yeah. So it maybe is a little bit easily.

Speaker B: So maybe the approach also can be not taking it that
seriously, be more playful and open to experiment, right?

Speaker A: Yes. Experiment, yeah yeah yeah. The way to experiment, to to
try, to to do the the to to find the right way to to to play music or
play art.

Speaker B: I see. So the final question, if you would define like your
life before your social media grow that big and after what is who you
were before and who you became after? Is there any difference?

Speaker A: I think the I am the same person but I am a lot of more sure
by myself because I I am

Notice the filler words, the false starts, the repeated "you you you should" — none of that was cleaned up. This is what the model actually heard, attributed to the actual speaker.

The SRT

Subtitle-ready output with millisecond timecodes, generated from the same diarized run. First five cues:

1
00:00:03,300 --> 00:00:04,200
I think that

2
00:00:06,800 --> 00:00:08,000
a really big problem

3
00:00:08,800 --> 00:00:11,800
and this is the voice because the voice is

4
00:00:11,800 --> 00:00:13,800
a very personal thing

5
00:00:14,600 --> 00:00:16,600
that you expose to the audience.

The VTT

Same segments, WebVTT format — note the WEBVTT header and the . decimal separator instead of SRT's ,, which is the actual difference between the two formats:

WEBVTT

00:00:03.300 --> 00:00:04.200
I think that

00:00:06.800 --> 00:00:08.000
a really big problem

00:00:08.800 --> 00:00:11.800
and this is the voice because the voice is

The PDF

The downloadable PDF transcript, first page, exactly as generated — same speaker labels, same text, formatted for reading rather than for a video editor:

First page of a convert.express PDF transcript, showing Speaker A and Speaker B labels on a diarized conversation about performing arts and social media

The AI summary

Generated by Claude from the finished transcript, in one pass, no edits:

The conversation focuses on the personal and vulnerable nature of using
one's voice as a singer, and how engaging with social media can serve
as valuable training for public exposure. Speaker A suggests that social
media is a good experimental space for artists, as it allows them to
practice communication and self-expression with the added safety net of
being able to delete posts, unlike live performances. Both speakers
agree that approaching social media playfully and experimentally, rather
than too seriously, is a healthy mindset for artists. Finally, Speaker A
reflects on personal growth through social media, noting that while
their core identity has remained the same, they have gained
significantly more self-confidence as a result of their online
presence.

The action items

Also generated by Claude, also unedited — and, honestly, empty:

No action items identified.

This is a reflective two-person conversation about art and social media, not a meeting with decisions or tasks. There was nothing to extract, and the model didn't invent anything to fill the space. That restraint is more useful to show than a fabricated task list would have been — on a real work call, this is exactly the output you'd want, minus the empty result.

The translated SRT

The same five cues, translated to Spanish, with the identical timecodes as the English SRT above:

1
00:00:03,300 --> 00:00:04,200
Creo que

2
00:00:06,800 --> 00:00:08,000
un problema realmente grande

3
00:00:08,800 --> 00:00:11,800
y esto es la voz porque la voz es

4
00:00:11,800 --> 00:00:13,800
algo muy personal

5
00:00:14,600 --> 00:00:16,600
que expones ante el público.

Compare 00:00:03,300 --> 00:00:04,200 on both versions — same numbers, same order, same cue count. That's not a coincidence and it's not extra engineering effort: translation reuses the stored subtitle segments from the original transcription pass rather than re-transcribing the audio in Spanish. The timing was already correct the first time; translation only ever touches the text field. Drop this file into the same video that used the English SRT and it lines up frame for frame, no retiming.

What this took

Transcription (241s of audio → diarized transcript):  ~17 seconds
Summary + action items:                                ~15 seconds
Spanish SRT translation:                                ~6 seconds
Minutes billed:                                         4.02 of 10 free daily minutes
Cost:                                                    €0.00 (within free daily quota)

Outside the free quota, 4.02 minutes at €0.02/minute would round up to convert.express's €0.50 minimum charge per transaction — summary, action items, and translation cost nothing extra either way, on top of that.

What convert.express does not do

To be direct about the edges of this: there is no player-synced editor where you drag subtitle cues against a video timeline. There is no visual subtitle timeline at all — SRT and VTT come out as files, not as an editing surface. There is no way to relabel "Speaker A" to an actual name once and have it apply everywhere; each job's speaker labels are its own. There are no team folders or shared workspaces — every job belongs to the person or anonymous session that created it. And there are no meeting bots that join a call automatically; you upload a recording after the fact.

If you need any of those, convert.express is the wrong tool. What it does — turn a recording into an accurate, diarized, multi-format transcript with an AI summary and correctly-timed translated subtitles, in about the time it took you to read this page — is everything above.