Vocal Analyzer Docs

← Back to app

Interpreting your first result

This page walks through a single, realistic analysis and explains what each tile means in two to three sentences. Every tile links into the fuller reference page for that metric. Read this end to end once, then use the reference pages as lookups.

Imagine you uploaded a 12-second clip of connected speech and the dashboard has finished rendering.

The pitch spectrum bar

The topmost bar in the Gender Perception Spectrum card shows your median pitch (F0) sitting as a white diamond on a three-zone strip. The three zones are the published central-tendency bands for perceived- masculine (blue), perceived-androgynous (gold), and perceived-feminine (pink) adult speech in English. If your diamond is halfway across the gold zone at roughly 160 Hz, that is the fastest read of where this recording lands — your pitch falls inside the overlap region the literature considers perceptually ambiguous.

Pitch is the strongest single cue to perceived gender in voice (Leung et al. 2018 found it accounts for roughly 42% of listener judgement variance), which is why it is the first bar. For the full story — the exact band boundaries, percentile statistics, and algorithm — see Pitch.

The resonance (formant dispersion) bar

The second bar in the same card shows formant dispersion, a whole- recording proxy for your apparent vocal-tract size. Lower dispersion corresponds to a longer apparent vocal tract (lower, more "masculine" resonance); higher dispersion corresponds to a shorter one. A diamond in the middle gold zone at ~1140 Hz means your resonance also sits in the overlap region.

Formant dispersion has weaker direct perceptual validation than speaking F0 (Pisanski & Rendall 2011), so the dispersion bar is a tracking-oriented overlay rather than a perceptual verdict. For the per-formant breakdown (F1, F2, F3), the per-vowel view, and the VTL estimate derived from dispersion, see Formants.

The voice-quality tile

Below the spectrum card, in the Voice Detail tile grid, a tile labeled something like "Voice Quality (H1-A3)" shows a dB value. The exact metric name depends on how much correction the analyzer could apply to your recording — the tool picks the most reliable rung of a preferred-metric ladder (H1-A3 → corrected H1-H2 → raw H1-H2). Higher values generally correspond to a breathier voice source; lower values to a pressed one.

At high F0 (above ~175 Hz) the spectral-tilt family loses gender- discrimination power, and the analyzer softens the claim accordingly. For the full ladder, the corrections, CPPS, and the LTAS complements (alpha ratio, Hammarberg index, spectral slope), see Voice quality.

The HNR tile

Another tile shows HNR (Voice Clarity) as a dB value. Higher numbers mean a cleaner, more periodic signal; lower numbers mean more noise or roughness in the source. Unlike pitch and resonance, HNR does not carry a gender band overlay — the published sex difference is only ~0–3 dB (Goy et al. 2013), too small to anchor a useful perceptual band. Treat it as a voice-health and recording- quality cue. See HNR.

Sibilants

If the recording has enough /s/-like frication, a Sibilant CoG tile shows the heuristic center-of-gravity frequency of the detected sibilants in Hz, along with spread / skewness / kurtosis details lower down. Higher CoG values correspond to a more concentrated, "sharper" sibilant spectrum. The detector is a heuristic — not phone-level segmentation — and is sensitive to the recording chain, so the dashboard pins sibilant confidence at moderate even in ideal conditions. See Sibilants.

Coherence

Between the primary spectrum bars and the secondary tiles, a Coherence tile shows whether your pitch percentile and your spectral-tilt percentile are telling the same story. Matched means both cues point the same way within the active profile; divergent means they disagree. This is a descriptive correlation cue, not a verdict. See Coherence.

A single measurement warning

Near the top of the summary panel, a small warning tile might say something like:

Recording appears noisy (SNR ≈ 14 dB). Pitch usually remains the most robust metric. Whole-recording formant and resonance estimates may still be usable, but interpret them with added caution under noisy conditions. Voice-quality and HNR values are the most likely to be biased.

That warning is informational: the tool still produced a full result, it is just flagging that the voice-quality and HNR values in particular should be read with added caution. The full enumeration of warning codes lives in Methodology.

What to do with this result

References