Your voice, without the room it was recorded in.
Upscale Voice splits the audio in your video into three layers — the speech, the steady background, and the narrow tones that sit on top — then lets you rebalance them by ear until the voice is the thing people hear.
What Upscale Voice does
Open a clip, and the audio is separated into speech, steady background, and narrow tones. From there you mix: turn the room down, bring the voice up, and hold Compare to check your edit against the original at any moment. When it sounds right, the enhanced audio is written back into the video.
- Fans, air conditioning, hum, traffic and other steady room noise
- Beeps, alarms, birdsong and phone notifications
- Speech that sits too low, too dull, or too uneven
How you mix it
Three faders
Voice, Background and Tones each get their own fader — turn the room down, bring the voice up.
Tone and Loudness
Tone runs warm to bright; Loudness lifts the level without crushing it.
Hold to Compare
Hold Compare at any moment to hear the original against your edit.
Solo a layer
Solo any layer to hear exactly what landed in it.
Go further
Pitch
Up or down 12 semitones, with length and lip sync preserved.
Space
Room, hall and plate reverbs.
Echo
Echo with adjustable time.
Spatial placement
Place the voice in the stereo field for headphones.
Character
Radio, phone, megaphone, robot and more.
Back into your video
The enhanced audio is written back into the video and saved to your photo library, ready to post.
Everything happens on your phone
No account, no sign-up, no upload. Your video never leaves the device, the app works in airplane mode, and nothing about you is collected.
Support & privacy policyNothing is uploaded
Your video is processed on the device. It is never sent to a server, ours or anyone else's.
No account required
There is no sign-up, no login and no profile. You open the app and start working.
Works offline
The whole workflow runs in airplane mode, because it never needed a connection to begin with.
- Works on the first 30 seconds of a clip.
- This is classic signal processing, not AI source separation.
- It handles fans, hum and room tone well — it can't isolate music or pull apart two people talking at once.
Does my video get uploaded anywhere?
No. All processing happens on your iPhone. The app has no account system and works with no connection at all.
How long a clip can I process?
Upscale Voice works on the first 30 seconds of a clip.
Can it separate two people talking at once?
No. This is classic signal processing rather than AI source separation. It is very good at steady noise — fans, hum, air conditioning, room tone — but it cannot isolate music or pull apart overlapping speakers.
What do I get at the end?
The enhanced audio is written back into your video and saved to your photo library, ready to post.
Is it available for Android?
Upscale Voice is built for iPhone / iOS.
