Vocal Analyzer Docs

← Back to app

Studio walkthrough

Studio is the surface where you read a single recording in depth. It is reached from the Studio link in the top navigation (the hash URL is #analyze), or by clicking Open in Studio on any recording from Home. This page tours it in the order you normally read it. For the narrative version see Interpreting your first result; for per-metric depth see the Metric reference sidebar section.

The web app is a single page with four surfaces — Home, Studio, Practice, and Settings — switched from the top navigation bar. Studio replaces what earlier versions called the "dashboard".

The top navigation

Every surface shares a sticky top rail:

There is no account or avatar — the app has no login and stores everything locally. See Storage and privacy.

The library rail

The left column of Studio is the Library: the master list of every saved session, newest first. Each row shows the recording name (or "Live session" for a live take), and a meta line like Jun 7 · 0:42 · median 182 Hz. Click a row to load it into the detail pane; hover to reveal a delete (×). The header has a collapse chevron so you can widen the reading area, and — when you have at least two comparable takes — a Compare A/B button that suggests a pair to line up side by side.

Before you pick anything the detail pane shows a Studio empty state with an Upload audio file button. You can also drag a file anywhere on the page (a "Drop audio file here" overlay appears) — WAV, MP3, or AIFF. Upload streams live progress ("Uploading…", "Analyzing…") and then renders the result.

The player bar

Once a session loads, a transport bar appears:

Reading · this take

The headline block gives you the one-glance read:

Perception spectrum

To the right, a Gender coding · on/off switch controls this surface's display (a per-page override; see below). When on, one or two spectrum rows — Pitch (F0) and Resonance · formant dispersion — show a marker over masculine / androgynous / feminine zones with a "reference overlay, not a verdict" footnote. When off, the reference bands are withheld and only your own target band shows on the meters. See Pitch and Formants.

Resonance is reported as formant dispersion in hertz. Vocal-tract length in centimetres is only a derived display unit you can switch on in Settings — it is never a 0–100 score.

The timeline

A canvas plots your pitch (amber) and resonance / formant dispersion (mint) contours over the recording, on a real hertz axis, with your target bands shaded. A Phrase / Full control zooms between a moving phrase window and the whole recording. Click to seek; the player follows.

When a transcript is available (installed app only), a synced transcript strip renders below the timeline, with phrase-consistency chips highlighting how steady each phrase was.

Metric patch-bays

Below the timeline, secondary metrics are grouped into family patch-bays in a fixed order: Pitch → Resonance → Weight & quality → Spectral & sibilants. A short legend, "How to read the evidence dots", explains the coloured dot on each cell:

These describe benchmark support for the measurement, not whether a recording will be perceived a certain way.

Each cell shows a short label, its evidence dot, an toggle that expands a plain-language explanation, the formatted value with unit, an optional mini target meter, and a status caption ("19 Hz below target", "within your expressiveness target", "derived from dispersion"). Jitter and shimmer fold into one cell; distribution and coverage metrics render as compact multi-part values.

Some cells carry a small confidence note when this recording's measurement quality is Moderate, Low, or Unavailable — that describes this recording, whereas the evidence dot describes the metric's broader support.

Vowel panels

When the analysis found vowels of enough distinct classes, two panels render from the per-vowel data. Vowels are found from the recording's sound (intensity peaks, voicing and stable formants), so these panels work without transcription, including in the browser version. In the installed app, aligned vowel intervals are used instead when phone alignment is available.

These hide entirely when a recording has no vowel data, and a mid-render failure clears them rather than showing the previous take's chart.

Benchmark continuity is anchored to whole-recording formant mode; the per-vowel view is additive. See Methodology.

Measurement warnings

When the extraction is uncertain, a status line collects per-section warnings (for example "sample rate below 16 kHz", "high-F0 Burg instability", "source-filter coupling", "low SNR", "insufficient usable formants"). The full list and plain-language explanations live in Methodology.

Comparing takes