Beyond the tool and further reading
Vocal Analyzer is one input into voice training, not the whole of it. This page points at the other pieces of the picture — the humans and the research — that an acoustic dashboard can complement but not replace.
A voice teacher is not replaceable by this tool
An acoustic dashboard can measure what came out of your mouth. It cannot hear what a voice teacher hears:
- Muscle tension you are holding without realising it.
- Breath support patterns that are setting you up for vocal fatigue.
- Resonance shaping (forward focus, nasality, larynx height) that happens in real time and interacts with every other cue.
- The difference between a technique that will sustain long-term and a compensation that will cause problems in six months.
If you are doing active voice feminization, masculinization, or androgynization work, finding a voice teacher experienced in that area will almost always be higher-leverage than any amount of acoustic measurement. Vocal Analyzer is useful alongside that work — tracking progress numerically, verifying that subjective changes match acoustic ones, catching regressions — not instead of it.
Community pointers
The voice-training community around feminization, masculinization, and androgynization has grown substantially since the 2010s and continues to expand. Rather than point at specific communities that may change or close, the most durable advice is:
- Look for resources that explicitly separate what acoustic cues are from what to practice — good resources will do the former themselves and defer the latter to a teacher.
- Be cautious of resources that promise a specific target pitch or formant "for" any given gender. The research is descriptive, not prescriptive (see Voice and gender perception).
- Check that any exercise or technique has been recommended by a teacher or SLP you trust before adopting it. "Heard it on the internet" is not the same as "vetted".
Research reading list
The reference bands, algorithms, and methodology decisions in Vocal Analyzer are grounded in published acoustic and perceptual research. Each per-metric reference page lists the specific citations that back its bands and algorithms. A compiled list:
Pitch and gender perception
- Leung, Y., Oates, J., & Chan, S. P. (2018). Voice, articulation, and prosody contribute to listener perceptions of speaker gender. Journal of Speech, Language, and Hearing Research 61(2), 266–280.
- Gelfer, M. P., & Bennett, Q. E. (2013). Speaking fundamental frequency and vowel formant frequencies: effects on perception of gender. Journal of Voice 27(5), 556–566.
- Davies, S., Papp, V. G., & Antoni, C. (2015). Voice and communication change for gender nonconforming individuals: giving voice to the person inside. International Journal of Transgenderism 16(3), 117–159.
- Holmberg, E. B., Oates, J., Dacakis, G., & Grant, C. (2010). Journal of Voice 24(5), 511–522.
- Pépiot, E. (2014). Male and female speech. Speech Prosody 2014.
Resonance, formants, and vocal-tract length
- Pisanski, K., & Rendall, D. (2011). The prioritization of voice fundamental frequency or formants in listeners' assessments of speaker size, masculinity, and attractiveness. The Journal of the Acoustical Society of America 129(4), 2201–2212.
- Reby, D., & McComb, K. (2003). Anatomical constraints generate honesty: acoustic cues to age and weight in the roars of red deer stags. Animal Behaviour 65(3), 519–530.
- Fitch, W. T. (1997). Vocal tract length and formant frequency dispersion correlate with body size in rhesus macaques. JASA 102(2), 1213–1222.
- Hillenbrand, J., Getty, L. A., Clark, M. J., & Wheeler, K. (1995). Acoustic characteristics of American English vowels. JASA 97(5), 3099–3111.
Voice quality, spectral tilt, CPPS
- Klatt, D. H., & Klatt, L. C. (1990). Analysis, synthesis, and perception of voice quality variations among female and male talkers. JASA 87(2), 820–857.
- Iseli, M., & Alwan, A. (2004). An improved correction formula for the estimation of harmonic magnitudes and its application to open quotient estimation. ICASSP 2004.
- Simpson, A. P. (2012). The first and second harmonics should not be used to measure breathiness in male and female voices. Journal of Phonetics 40(3), 477–490.
- Chai, X., & Garellek, M. (2022). On H1–H2 as an acoustic measure of linguistic phonation type. JASA 152(3).
- Hillenbrand, J., & Houde, R. A. (1996). Acoustic correlates of breathy vocal quality: dysphonic voices and continuous speech. Journal of Speech and Hearing Research 39(2), 311–321.
- Heman-Ackah, Y. D., Michael, D. D., & Goding, G. S. (2003). Journal of Voice 17(1), 20–27.
Algorithmic background
- De Looze, C., & Hirst, D. J. (2008). Detecting changes in key and range for the automatic modelling and coding of intonation. Speech Prosody 2008.
- Boersma, P. (1993). Accurate short-term analysis of the fundamental frequency and the harmonics-to-noise ratio of a sampled sound. Proceedings of the Institute of Phonetic Sciences 17, 97–110.
Perception, clinical, and normative reference data
- Dacakis, G., Oates, J. M., & Douglas, J. M. (2017). Beyond voice: perceptions of gender in male-to-female transsexuals. International Journal of Transgenderism 18(1), 43–54.
- Goy, H., Fernandes, D. N., Pichora-Fuller, M. K., & van Lieshout, P. (2013). Normative voice data for younger and older adults. Journal of Voice 27(5), 545–555.
- Awan, S. N., et al. (2024). Bridge2AI Voice Technical Report.
Speaker embeddings (optional pipeline)
- Wang, H., et al. (2023). WeSpeaker: a research and production oriented speaker embedding learning toolkit. ICASSP 2023.
- Snyder, D., et al. (2018). X-vectors: robust DNN embeddings for speaker recognition. ICASSP 2018.
Product history
For a chronological record of what changed across Vocal Analyzer
releases, see CHANGELOG.md in the repository. The changelog is
written from a developer perspective but is the canonical record of
which metrics appeared when and which heuristics changed.