Accuracy and limitations

Accurate enough to inspect; transparent about what is not proven

VoiceScope includes regression tests for octave correction and Smart filtering, but it is not a calibrated medical instrument and has not completed a formal multi-device laboratory benchmark.

Three evidence levels separating regression tests, real-world repeatability, and unclaimed clinical validation
Algorithm behaviour, repeatability, and clinical validation are different evidence levels.

What is currently tested

  • Short bracketed octave mistakes are corrected without dragging later speech into the wrong octave.
  • Brief low-confidence local spikes can be removed.
  • Clear short phrases and sustained register changes are preserved.
  • Mobile layout regressions and localization completeness have automated checks.
Repeatable voice recording setup with stable microphone, distance, and duration
Matched recording conditions reduce variation that does not come from the detector itself.

Likely sources of error

Octave errors
A harmonic may be mistaken for the fundamental or vice versa.
Noise
Fans, music, other voices, clicks, and reverberation can produce false candidates.
Vocal quality
Breathy, creaky, rough, or rapidly changing phonation can reduce detector confidence.
Hardware
Microphones and browser audio paths vary by device.

How to check repeatability

Use the same device, room, passage, distance, and speaking task. Record three samples and compare their medians and contours. A result that changes dramatically under matched conditions deserves inspection rather than averaging without explanation.

What VoiceScope does not claim

The tool does not diagnose vocal disorders, determine gender, score attractiveness, or estimate a complete singing range from ordinary speech. It also does not claim to be more accurate than another product without a shared reference dataset.

Planned benchmark

A useful next validation step is a public set of clean tones, synthetic speech-like signals, and labelled human-voice recordings with expected F0. Results should report gross pitch error, octave error rate, voiced/unvoiced mistakes, and device repeatability. Until that exists, the documented algorithm and inspectable contour are evidence of transparency—not clinical validation.

🇬🇧 English 🇷🇺 Русский 🇪🇸 Español 🇩🇪 Deutsch 🇫🇷 Français 🇧🇷 Português 🇨🇳 中文 🇯🇵 日本語 🇰🇷 한국어 🇮🇳 हिन्दी