VoiceSwap.AIHow It Works
The full pipeline from raw video to download — all automatic.
Upload Video
Drop any MP4 or MOV. Up to 2 minutes.
Gemini Analyzes
Transcription, emotion, tone and pacing in a single API call.
Pick a Voice
6 Chirp3 HD voices — the highest quality Google TTS offers.
TTS Synthesizes
Gemini-authored SSML drives a truly expressive performance.
Download
New voice merged into your video. Ready to share.
Why VoiceSwap
Not just a transcriber. Gemini detects emotion segments, emphasis words, and writes expressive SSML performance notes — the creative intelligence of the pipeline.
Google Cloud TTS receives Gemini-authored SSML so it speaks with the same energy, pace, and emotional emphasis as the original speaker. Different voice, same soul.
Upload → Gemini analyze → SSML → TTS synthesize → ffmpeg merge. The entire pipeline runs automatically and delivers a download-ready video in under a minute.
The Intelligence
Transcribes
Full text with word-level timestamps.
Analyzes emotion
Tone, energy, pace, and emphasis per segment.
Writes SSML
Tells Google TTS exactly how to perform the text.
{
"overall_tone": "calm and explanatory",
"pace": "moderate",
"energy": "medium",
"emphasis_words": [
"important", "remember", "key"
],
"ssml_hints": "Speak clearly with slight enthusiasm, pause naturally after key points"
}Generated by Gemini 2.5 Flash · Drives SSML for Google TTS