Silence Remover

Long pauses in a spoken recording are found automatically — anything quieter than the level you set — and then either shortened or cut out altogether. It is the quickest way to tighten up a talk, a podcast or a voice note.

How to use it

The two settings that matter are the threshold, which decides how quiet counts as silence, and the minimum length, which stops the natural gaps between words from being clipped.

  1. Open the recording The waveform shows you what you are working with. The flat parts are the gaps.
  2. Set the threshold just above the room noise Start at minus 45 dB. If it cuts into speech, go lower; if it misses gaps, go higher.
  3. Apply and listen to the joins Shortening gaps rather than removing them keeps the rhythm of speech. Removing every pause sounds unnatural fast.
Full instructions

Drop an audio file here

or

Your file is opened on this device. It is never uploaded.

Open a file to see its waveform here.

Silence

In depth

How to cut dead air from a podcast or voice recording

Long pauses in a spoken recording make it drag. This guide walks through every control on the Silence Remover page, shows how the threshold, gap length and padding settings work together, and gives a step-by-step example of tightening a ten-minute voice note down to seven. By the end you will know what each slider does, which output format to pick, and how to avoid the two mistakes that ruin an otherwise clean edit.

Screenshot: Taking the dead air out
Taking the dead air out as it appears when the page opens.

What the tool does in one sentence

It scans the full waveform of your recording, finds every stretch that sits below a volume level you choose and lasts long enough to count, and then shortens or removes those stretches. The result is a tighter file with the same speech, the same pitch and the same quality, just fewer empty spots between the words.

Podcasters, teachers, students reviewing lecture captures, and anyone who records voice memos on a phone all run into the same problem. A five-second pause while you check your notes is fine in the room but painful on playback. Cutting each pause by hand in an editor is accurate but slow. This page does the same job in a few seconds, with controls that let you fine-tune how aggressively it cuts.

Opening a file and reading the waveform

Drop an audio file onto the page, or press the button to browse for one. The formats it accepts include MP3, WAV, M4A, AAC, OGG, OPUS, FLAC, and WebM audio. Once the file loads, a waveform appears. The flat, low sections are the gaps. The tall spikes are the loud parts - speech, music, or sound effects.

Below the waveform you will see two time stamps and a selection length. You can drag across the wave to pick a region, or press Select all to work on the whole recording. A Play button lets you listen before and after any change, and a zoom indicator shows how far in you are.

The six controls that shape the cut

All the settings sit in the panel labelled Silence. Each one has a slider and a live readout so you can see the value as you drag.

  • Anything quieter than sets the threshold. It runs from minus 70 dB to minus 20 dB, with the default at minus 45 dB. Any audio that stays below this level counts as silence. If the tool misses obvious gaps, raise this value. If it clips into speech, lower it.
  • and longer than sets the minimum gap length. It runs from 100 ms to 3000 ms, with the default at 500 ms. A short breath pause of 200 ms will not be touched at the default setting. Raise this to skip small gaps and only target the long dead spots.
  • Do this to it offers three choices. Shorten it to the length below keeps a little pause in place of the original gap. Remove it entirely cuts every gap out. Leave it alone - just show me where it is marks the gaps on the waveform without changing anything, which is useful for a first pass.
  • Shorten gaps to controls how long each gap becomes after the cut. The range is 0 to 1000 ms, and the default is 200 ms. Setting it to zero is the same as removing the gap altogether.
  • Leave a little either side adds padding before and after each cut. The range is 0 to 300 ms, and the default is 40 ms. This protects the tail end of one word and the start of the next from being clipped.
  • Crossfade the joins blends the audio at each edit point. The range is 0 to 60 ms, with 12 ms as the default. Without a crossfade, each join can produce a tiny click. Even 10 ms is enough to smooth it out.

The trim checkbox and the Apply button

Below the sliders is a checkbox labelled Also trim the start and the end. It is ticked by default. When active, any silence at the very beginning and very end of the file is cut in the same pass. Untick it if your recording needs a clean lead-in or tail.

Press Apply to run the detection. The waveform redraws to show the result, and a status line tells you how many gaps were found, how many seconds were removed, and what percentage of the total that represents. Press Start again at any time to undo everything and return to the original waveform.

Choosing a format and saving the result

The bottom panel has three controls for the output file. Save as lists the formats your browser can encode. M4A and WebM are compressed choices that keep file sizes small. WAV is lossless but much larger. MP3 is an option. Browsers can play MP3 files but cannot write them, so this page loads a small open-source encoder from this site the first time you choose MP3. Your audio is still never uploaded anywhere.

WAV quality offers 16-bit for normal use, 24-bit for studio work, and 32-bit float for further editing. Compressed quality sets the bitrate for M4A or WebM, from 32 to 320 kbps, with 128 as the default. Press Save the tightened audio and the processed file downloads to your device. A progress bar tracks the encoding.

A worked example with real numbers

Say you have a ten-minute interview recorded on a phone. The room is quiet but there is a faint hum from an air conditioner. Drop the file in. The waveform shows clear peaks where someone speaks and flat valleys between answers.

Start with the defaults: threshold at minus 45 dB, minimum gap at 500 ms, shorten mode with gaps kept at 200 ms, 40 ms of padding, and 12 ms of crossfade. Press Apply. The status bar says it found 23 gaps totalling 2 minutes and 48 seconds, or about 28 percent of the file.

Play the result. The speech flows well, but one join sounds rushed because the speaker took a deliberate breath before a long answer. Raise the minimum gap from 500 ms to 800 ms and press Apply again. Now only the truly dead spots are caught, and the natural pauses remain. The total cut drops to 2 minutes and 12 seconds. That feels right. Set Save as to M4A and press Save the tightened audio. The file downloads at about 1 MB per minute of audio.

What to do when the threshold seems wrong

The most common problem is that gaps are not found at all. This happens when the background noise in the room is louder than the threshold. An air conditioner, a computer fan, or street noise can push the floor above minus 45 dB. The fix is to raise the threshold toward minus 30 dB until the tool starts catching the gaps. Move in steps of 5 dB and check after each one.

The opposite problem is that speech gets clipped. This usually means the threshold is too high, or the padding is too low. Lower the threshold by 5 dB and raise the padding from 40 ms to 80 ms. Between the two changes the ends of words should survive. Recordings with a very high noise floor sometimes have no threshold that works, because the room is audible throughout and there is never a true gap to find.

Shorten versus remove, and why it matters

Removing all gaps completely can make speech sound frantic. A person who pauses between thoughts sounds natural. A person with every pause deleted sounds like they are rushing through a script. Shortening each gap to 200 ms keeps a small beat between phrases. The listener still feels the rhythm of the original recording, even though two or three minutes of dead air are gone.

The mark option is worth trying before you commit to any cut. It tells you how many gaps your current settings would find and how much time they add up to, without changing a single sample. If the count is higher than you expected, adjust the minimum gap length before you apply.

When this tool is not the right choice

If the background noise is the problem rather than the silence, this page will not help. It only finds quiet stretches and cuts them. It does not reduce hiss, hum, or fan noise. A noise reduction tool is what you need in that case.

If the audio has music underneath the speech, the music fills the gaps with sound and the detector will not find anything to cut. The same goes for recordings with two or more speakers who overlap regularly. In both cases, manual editing in a full audio editor is the more reliable path.

Your recording stays on your machine

The entire process runs inside your browser. The file you drop in is decoded locally, the gaps are found locally, and the output is encoded and saved locally. Nothing is sent to a server at any point. You can test this yourself by loading the page, opening a file, turning off your internet connection, and running the full edit. It works exactly the same way.

That privacy matters here more than it might for other tools. Voice recordings often contain names, phone numbers, business discussions, or personal conversations. Knowing that the audio never left your device means there is nothing to worry about on that front.

Help

Taking the dead air out

The two settings that matter are the threshold, which decides how quiet counts as silence, and the minimum length, which stops the natural gaps between words from being clipped. Every stretch that meets both is then either shortened to a set length or taken out. A room with a hum in it needs a higher threshold than a treated one.

Open the recording

The waveform shows you what you are working with. The flat parts are the gaps.

Set the threshold just above the room noise

Start at minus 45 dB. If it cuts into speech, go lower; if it misses gaps, go higher.

Apply and listen to the joins

Shortening gaps rather than removing them keeps the rhythm of speech. Removing every pause sounds unnatural fast.

Silence is rarely actually silent

A recording made in a real room has a noise floor — air conditioning, a computer fan, the room itself — and that noise is what the threshold is really measuring against. If your gaps are not being found, the noise floor is above your threshold, and raising it is the fix. If speech is being cut, the quiet ends of words are falling below it, and lowering it is the fix. A recording with a high noise floor may have no threshold that works, in which case there is no gap to find: the room is audible throughout.

Good to know

Shorten rather than remove

Speech has a natural rhythm. Cutting every pause to nothing makes a person sound like they are being chased.

A crossfade hides the joins

Ten milliseconds is enough to stop each cut making a faint click.

Try the mark option first

It tells you how many gaps your settings would find, before anything is changed.

Common questions

How do I remove silence from a recording automatically?
Open the file, leave the defaults, press Apply, and check the joins. Adjust the threshold if too much or too little was cut.
It cut off the ends of my words. What now?
Lower the threshold, and raise the amount left either side of each gap.
It did not find any gaps.
The background noise in the recording is louder than the threshold. Raise the threshold towards minus 30 dB.
How much time will this save me?
The result tells you exactly how much was cut, in seconds and as a percentage.
Is my recording uploaded?
No. It never leaves your device.