Audiobook Mastering Studio

Check narration against the ACX submission requirements and master it to pass, entirely inside your browser. Free, fast, and your files never leave your device.

🎙️

Drop your chapter files here or click to choose

WAV, MP3, FLAC, M4A and OGG. Add as many chapters as you like.

Your recordings never leave this device. Every file is decoded, measured, mastered and encoded by your own browser. Nothing is uploaded, nothing is stored, and no account is needed.
Advanced: targets and noise gate
  • RMS level: -23 to -18 dBFS (mastered to -20)
  • Peak level: -3 dBFS or lower (limited to -3.5)
  • Noise floor: -60 dBFS RMS or lower (targeted at -65)
  • File length: under 120 minutes
  • Room tone at head: 0.5 to 1.0 s (padded to 0.75)
  • Room tone at tail: 1.0 to 5.0 s (padded to 2.5)
  • Export: MP3 192 kbps CBR, 44100 Hz
Stronger settings remove more room noise but can shorten breaths and soft word endings. Start with Standard and only move up if the noise floor still fails.

ZillaKit is not affiliated with, endorsed by, or connected to ACX, Audible or Amazon. The name ACX is used here only to describe the published submission requirements that this tool measures against. Always confirm the current requirements with the platform before you submit.

How to use the Audiobook Mastering Studio

  1. Drop your narration files into the box above, one file per chapter. WAV, MP3, FLAC, M4A and OGG all work.
  2. Click Check all for a free ACX audio requirements check. You get measured RMS, peak, noise floor, length and room tone for every chapter, with a pass or fail badge against each target.
  3. Click Master all to run the full mastering chain: rumble filter, noise gate, RMS normalisation to -20 dBFS, peak limiting and room tone padding.
  4. Open Details on any row to see the before and after compliance table with every measurement side by side.
  5. Download each mastered MP3, or use Download all + report to get a zip containing every MP3 plus a plain text compliance report.

Why use ZillaKit's audiobook compliance check?

Most audiobook rejections come down to four numbers, and all four are measurable before you ever submit. Your narration has to sit between -23 and -18 dBFS RMS, peak no higher than -3 dBFS, keep a noise floor at or below -60 dBFS RMS, and carry a short stretch of room tone at the head and tail of every file. Miss any one of them and the whole chapter comes back. This tool measures all of them from the decoded audio, then fixes what can honestly be fixed, so you can master an audiobook for ACX online without buying a digital audio workstation or learning a compressor.

The order of operations is the part that trips people up. RMS normalisation is just a gain change, so raising a quiet recording to -20 dBFS raises the room noise by exactly the same amount. A file that measured a -55 dBFS noise floor at -30 dBFS RMS will measure a -45 dBFS noise floor once it is normalised, and it will fail. That is why this tool measures the noise floor first, applies a downward expander sized to the gain that normalisation is about to add, and only then normalises. It is the correct ACX RMS noise floor fix, and it is why free audiobook mastering that only does volume adjustment leaves so many narrators confused about why their files keep bouncing.

Everything runs locally. Your browser decodes the audio, processes the samples, encodes the MP3 with a WebAssembly encoder and hands you the file. Unreleased manuscripts, client work under contract and personal recordings never touch a server, which matters when the material is not yours to share. There is no signup, no watermark, no email capture and no queue, and because the work happens on your own device the speed depends on your machine rather than someone else's load. The compliance report is plain text, so you can keep it with the project files or send it to a rights holder as evidence of what was measured.

What the checker measures

RMS is measured over 400 millisecond windows with the quietest ten percent discarded, so long lead-in silence and pauses between paragraphs do not drag the average down and give you a falsely quiet reading. Peak is a true sample peak taken across every sample of every channel, not an estimate. The noise floor is the RMS of the quietest 400 millisecond window that is not digital silence, which is the same idea a mastering engineer uses when they find the gap between sentences and read the meter. Room tone at the head and tail is measured as the run of audio below a speech threshold derived from the file's own level and floor, so it adapts to loud and quiet recordings alike.

FAQ

What are the ACX audio requirements?

Each file must measure between -23 and -18 dBFS RMS, peak at -3 dBFS or lower, have a noise floor at or below -60 dBFS RMS, run under 120 minutes, and open and close with room tone rather than digital silence, roughly 0.5 to 1 second at the head and 1 to 5 seconds at the tail. Submitted files are 192 kbps or higher constant bit rate MP3 at 44.1 kHz, mono or stereo, and consistent across the whole book. This tool measures against those published targets and exports 192 kbps CBR MP3 at 44100 Hz.

Why is the limiter ceiling set to -3.5 dB instead of -3 dB?

Compliance is measured on the encoded MP3, not on the audio you fed the encoder. Lossy encoding reconstructs the waveform from frequency data, and the reconstructed samples routinely overshoot the original by a few tenths of a decibel, sometimes more on sibilance. Limiting to exactly -3.0 dBFS before encoding therefore produces a file that reads slightly above -3.0 dBFS afterwards and fails. Half a decibel of headroom absorbs that overshoot while staying comfortably inside the acceptable range, so the delivered MP3 passes.

Why can a noisy recording sometimes not be fixed automatically?

A downward expander can only lower the parts of the file that are already quiet. It cannot remove hiss, traffic, air conditioning or computer fan noise from underneath the narration itself, because that noise is mixed into the same samples as your voice. If the gap between your voice and the room noise is small, there is nothing to gate: any setting strong enough to hit -60 dBFS would also chew through breaths, soft consonants and word endings. When that happens this tool says so plainly instead of marking the file as passed. The real fix is at the source: a quieter space, soft furnishings, a closer microphone position and a recording level that puts your voice well above the room.

Do my files get uploaded anywhere?

No. Decoding, analysis, mastering and MP3 encoding all happen inside this browser tab using the Web Audio API and a WebAssembly encoder. No audio, no measurements and no filenames are sent anywhere. You can disconnect from the network after the page loads and the tool still works.

Should I export mono or stereo?

The tool keeps whatever the source has: a mono recording exports mono, a stereo recording exports stereo. What matters is consistency, since every file in a single book has to match. Solo narration is normally recorded and delivered in mono, which also halves the file size at the same bit rate, so if some chapters are mono and others are stereo, convert everything to one format before you submit.

How long a file can I process?

Chapter length files of roughly 10 to 30 minutes are the safe zone and match how audiobooks are submitted anyway. Longer files need a desktop browser with plenty of free memory, because the entire recording is held as uncompressed samples while it is processed. If a very long file fails to process, split it into chapters and run them separately.

Is this tool affiliated with ACX or Audible?

No. ZillaKit is independent and has no affiliation with, endorsement from, or connection to ACX, Audible or Amazon. Those names appear only to describe the published submission requirements this tool measures against.