1. Audio capture
The browser requests a microphone stream with local background-noise suppression. Echo cancellation and automatic gain control remain disabled so they do not reshape the pitch signal. PCM samples are captured locally and downsampled to approximately 16 kHz for analysis.
2. Fundamental-frequency estimation
Overlapping frames are analysed with the YIN algorithm, a time-domain estimator based on a difference function and cumulative mean normalization. Each frame produces an F0 candidate and a confidence value. The original method is described by de Cheveigné and Kawahara in YIN, a fundamental frequency estimator for speech and music.
3. Acceptance and correction
Frames outside the configured frequency range or below detector confidence are rejected. A recording-level evidence gate rejects noise-only contours made from scattered YIN fallbacks or short mechanical sounds. An adaptive filter also removes quiet, persistent background-tone readings while preserving supported speech and stable notes. Local octave correction tests brief 2× or ½× errors against unchanged neighbours on both sides.
4. Smart filtering
Smart mode uses a hybrid median-absolute-deviation detector. The primary test compares a short run with stable context on both sides. A conservative fallback handles isolated extremes near pauses. Detector confidence protects clear pitch changes: low-confidence anomalies may be removed, while sustained runs survive.
5. Statistics and visualization
Accepted frequencies are converted to MIDI semitone positions and note names. The app calculates median, arithmetic mean, observed frequency range, full min–max semitone spread, and voiced proportion. Moving-average and exponential lines are visual aids; the processed detected points remain available.
The complete implementation is available in the source repository.