TRACKED VOCALS

A working definition

Noise floor: a working definition for session vocals.

Here is what we measure and why, so you can check our work. House rules live on The Standards. This page is the method — useful whether or not you ever list here.

The gap

What is specifiedWhat it says
LoudnessITU-R BS.1770-4, EBU R 128. Netflix Sound Mix Specifications v1.6 specifies -27 LKFS and -2 dBTP (source).
Listening roomsEBU Tech 3276 §2.6: background noise should preferably not exceed NR 10… under no circumstances should it exceed NR 15. Explicitly scoped to listening and monitoring, not capture.
MicrophonesSelf-noise is published in dB-A for every studio mic sold. Neumann: below 10 dB-A extremely low; 11–15 very good; 16–19 good enough; 20–23 high for a studio mic; 24+ unworthy of a studio microphone.
Vocal captureThe Recording Academy P&E Wing delivery recommendations — the closest thing to a music-industry delivery document — contain no noise-floor requirement. Netflix’s music-adjacent spec contains none. We could not find a published scale for session vocal capture.

So the singer standing between a specified microphone and a specified listening room, with nothing specified in between, reaches for the only number they can find — ACX’s -60 dB RMS, an audiobook spec, measured on a few seconds of hand-picked silence, from chains where printed compression isn’t permitted. It doesn’t fit their work, and it makes good recordings look like failures.

Neumann notes that “even a very quiet recording room will contribute quite a bit more ambient noise than 10 dB-A.” The room, not the microphone, is the binding constraint. That is why this method measures the room.

The method

Reproducible or it isn’t a definition. Someone else should be able to implement this section and get our number on the same file.

  1. Split the file into non-overlapping windows of 50 ms. Compute the mean-square (RMS energy) of each window, unweighted, across channels.
  2. Convert each window to dBFS: 10 · log10(mean-square). Digital silence is −∞.
  3. Sort the window energies ascending. Take the arithmetic mean of the quietest 10%. That mean, in dBFS, is the floor.
  4. Measure a dry pass (file floor) or an unprocessed room-tone capture of 5–10 s (room floor). Never a mixed reel — its quietest windows are still music.
  5. Assume ordinary session gain staging: peaks around -6 to -10 dBFS. Record hotter and the same room prints a lower number.

The same implementation runs in the browser Vocal Reel Certifier and on the server. The two cannot disagree about a file.

File floor vs. room floor

These are two quantities. Conflating them is the usual error.

File floor

what a producer receives — your printed compression included.

Quietest 10% of 50 ms windows on the delivered dry pass.

Room floor

from your room-tone capture — the space itself, unprocessed.

The same measurement on 5–10 seconds of silence, no processing.

Compression’s makeup gain raises the file floor dB-for-dB, so a clean room with printed compression reads 5–6 dB worse than it is. The spoken-word world never articulates this because printed compression isn’t allowed there — so the question never arises. In music, where it’s normal practice, it’s the crux. Report both or report neither.

The four bands

Names describe how audible the floor is — never what the file is for, and never the singer. A boundary without a derivation is an opinion.

BandRange (dBFS)Where the number came from
Studio silentbelow -70Professional studios typically sit at -80 to -70 dBFS.
Inaudible floor-70 to -58Well-treated home rooms typically sit at -70 to -65 dBFS.
Library quiet-48 to -58Untreated rooms typically sit at -50 to -40 dBFS. The -58 split is the least evidence-backed of the three — the published data leaves a gap between -50 and -65. It is under active calibration against delivered, client-approved session files.
Audible noiseabove -48Above the untreated-room cluster. The floor is audible beside the voice.

-58 is the least evidence-backed of the three ceilings. The published data leaves a gap between -50 and -65. That split is under active calibration against a corpus of delivered, client-approved session files.

What this measurement cannot do

A quietest-decile floor is a steady-noise measurement. It is blind to anything intermittent, and it understates tonal hum.

FailureCaughtWhy
Fans, AC, HVAC, computer noiseYesContinuous and broadband.
Mains hum / buzzPartiallyContinuous, but broadband RMS badly understates tonal noise.
Sustained intermittent noise — traffic, HVAC cyclingPartiallyLong enough to occupy a tenth of the quiet windows, so the dry-pass spread sees it.
Brief isolated events — a door, a chair creak, one barkRoom tone onlyOn a dry pass the flag compares the 90th percentile of gap windows to the 10th; a 200 ms event is ~4 windows in ~100 and never reaches the 90th. A room-tone capture uses the 98th and does catch it.
Concealed edits in gaps / silenceNoA cut placed inside silence leaves no discontinuity to measure, so a well-hidden edit in a gap is undetectable. Mid-phrase clicks and splices are caught; edits concealed in silence are not.
Radio / RF bleedNoModulated and speech-like.
A phone recordingNoCan post a clean floor; its faults are bandwidth, codec, AGC and distance.

The noise this measurement catches well is the easy kind to remove; the noise it misses is the hard kind. A file at -66.8 with mains hum and something walking through the gaps scores better than a clean -53 and would be sent back by any producer.

Clipping is in the Certified Reel seal. The floor is not. One is a defect a contractor cannot use; the other is a description of a room. A clipped reel still publishes.

The two flags

They exist because of section 5. They reuse the 50 ms windows. They do not gate publication and they do not gate certification.

Mains hum

Goertzel at 50 Hz and 60 Hz and their first two harmonics. A peak 12 dB above the local spectral median of surrounding bins flags. Both mains frequencies are tested — we do not assume 60. Broadband RMS badly understates a tone; this is the check that can see it.

Intermittency

On a dry pass, gaps are windows within 20 dB of the floor estimate — not a fixed fraction of the file. Compare the 90th percentile to the 10th. Breaths are normal and sit in the high tail; p90 leaves them out. On room tone every window is a gap and the high statistic is the 98th percentile, so a genuine event is visible and a single-sample click is not. A spread of 12 dB flags.

Both thresholds (12 dB and 12 dB) are provisional and being calibrated against real files.

Version

Method v1.0 — August 2026.

If you measure differently and get a different result, or you know of a published scale for session vocal capture we missed, tell us: studio@trackedvocals.com. A person reads that address.

Sources

Figures on this page are imported from the measurement module. If they ever disagree with the validator, the validator is right.