Vocal Analyzer Docs

← Back to app

Glossary

Short, plain-language definitions for the acoustic terms Vocal Analyzer uses on the dashboard and across these docs.

Burg method

An LPC variant Praat uses internally for formant estimation. The relevant fact here is that Burg-derived formants become less reliable at high F0 (see "high_f0_burg_instability" on the Methodology page).

Cepstrum / CPP / CPPS

The cepstrum is a transform of the spectrum that makes periodicity easy to measure. CPP (Cepstral Peak Prominence) and CPPS (smoothed variant) measure how prominent the main cepstral peak is — a high CPPS means a clean, periodic signal; low CPPS means a noisy or rough signal. See Voice quality.

F0 (fundamental frequency)

The rate at which your vocal folds vibrate during voiced speech, measured in Hz. Directly perceived as pitch. See Pitch.

F1 / F2 / F3 / F4

The first through fourth formants — acoustic resonances of the vocal tract. F1 varies with vowel height, F2 with frontness, F3 with rounding, F4 with overall tract length. See Formants.

Formant

See F1/F2/F3/F4 above.

Formant dispersion

The average spacing between successive formants, used as a whole- recording proxy for apparent vocal-tract size. Lower dispersion → longer apparent tract. See Formants.

HNR (harmonics-to-noise ratio)

The ratio of periodic (harmonic) to aperiodic (noise) energy in a voiced signal, in dB. Higher = cleaner voice. See HNR.

LPC (linear predictive coding)

A parametric model of the vocal-tract filter that lets the analyzer extract formant peaks from speech. Praat uses Burg-method LPC.

LTAS (long-term average spectrum)

The averaged frequency spectrum across the whole recording. The alpha ratio, Hammarberg index, and spectral slope tiles are LTAS- derived complements to the spectral-tilt ladder. See Voice quality.

Semitone

A 12th of an octave. Pitch variability is reported in semitones because a semitone represents the same perceptual step at any frequency — 5 Hz of variation at 100 Hz is perceptually much larger than 5 Hz at 250 Hz. See Pitch and Pépiot (2014).

Sibilant

A high-frequency fricative sound like /s/ or /ʃ/. Vocal Analyzer's sibilant detector is a heuristic, not phone-level segmentation. See Sibilants.

Source-filter model

The acoustic model where the vocal folds act as a periodic source and the vocal tract acts as a filter that shapes the source spectrum. Most of what listeners call "voice quality" lives in the source part of this model. See Voice quality.

Spectral tilt

How quickly energy in the source spectrum rolls off with frequency. Breathier sources have a steeper tilt; pressed sources have a shallower tilt. Measured by the H1-H2 family of metrics.

Spectrum

The frequency-domain representation of a signal: which frequencies are present and at what energy. The Fourier transform of the time- domain waveform.

Vocal tract length (VTL)

The effective acoustic length of the vocal tract from glottis to lips. Longer tracts shift formants down; shorter tracts shift them up. Vocal Analyzer estimates VTL in cm from formant dispersion via the Reby & McComb (2003) regression.

Voiced frame / voicing

An analysis frame where the vocal folds were vibrating (producing F0). The pitch tracker labels each frame as voiced or unvoiced based on periodicity and intensity.