Two Microphones or Four? Hardware Heterogeneity in In-Home Voice Measurement
Motivation
Not all microphones are equal, and the difference costs money. Our IHearYou array in a real home mixes two node types: the ReSpeaker Lite, which carries two microphones and does little with them, and the XVF3800, which carries four and combines them by beamforming, steering its sensitivity toward whoever is talking and suppressing the rest of the room. Beamforming is the expensive option, and the reasoning behind it is intuitive, since four microphones aimed at a speaker should hear that speaker more faithfully than two microphones aimed at nothing in particular.
Problem statement
Whether that intuition holds for what we actually measure is a separate question, and nobody has checked, including us. The array does not transcribe speech, it extracts vocal features such as pitch, jitter, and shimmer, which are fine-grained properties of the voice rather than gross properties of the recording. Beamforming improves intelligibility, which is what it was designed for, but it also filters the signal, and a filter that helps a listener understand a word can distort the very microtiming and amplitude detail a vocal feature depends on. In our own deployment the two node types sit in different rooms, so no heterogeneous pair has ever heard the same utterance, and without simultaneous capture the comparison cannot be made at all. Anyone building an in-home vocal sensing system is currently guessing at a purchasing decision.
The project
We will rearrange the array so that a Lite and an XVF3800 observe the same speech simultaneously, then score both against the truth. The truth comes from replay, since a known recording played into the room has feature values computable from the original file, so each node’s report can be measured as an error rather than merely compared to its neighbour. Automating that replay requires speaker playback on the node firmware, which does not exist yet and which this project will implement, and which incidentally unblocks three further experiments in the wider study. Reference speech comes from the TESS corpus, so the study needs no participants and touches no personal data. No background in embedded programming or acoustics is required, and the hardware is provided. The specific goal of this project is to determine whether four-microphone beamforming measures vocal features more accurately than two plain microphones, and to establish at what distance and under which noise conditions any advantage appears or disappears.
Objectives
- Understand the IHearYou signal chain from microphone to extracted vocal feature, and what beamforming does to the signal on the way.
- Implement calibrated speaker playback on the ESP32-S3 node firmware, triggered over MQTT, so that replay becomes automated and repeatable.
- Rearrange the array so that a ReSpeaker Lite and an XVF3800 capture the same utterance simultaneously.
- Measure per-feature error against the reference clips for both node types across distance (1, 2, 4, and 6 m) and noise condition (quiet, television, kitchen).
- Report where beamforming helps, where it does not, and derive a placement and purchasing guideline.
Previous theses in this line
- Audio-centered Approach for Building a Multimodal Predictive AI Agent to Detect Depressive Behaviors, Jonas Länzlinger.
- Algorithmic Framework for Sound Localization in Polyphonic Noisy Environments, Tibor Haller, Jonas Länzlinger.
References
[1] J. Länzlinger, K. O. E. Müller, B. Stiller, and B. Rodrigues. Towards Interpretable Depression Detection: Linking Acoustic Features to DSM-5 Indicators. IEEE PerCom 2026, Work-in-Progress.
[2] Khamaisi, K., Keller, N., Krummenacher, S., Huber, V., Fässler, B., & Rodrigues, B. (2025). From Noise to Knowledge: Acoustic Anomaly Detection in Pumped-storage Hydropower Plants. arXiv:2509.22881.
[3] Haller, T., & Länzlinger, J. (2025). Algorithmic Framework for Sound Localization in Polyphonic Noisy Environments. Interdisciplinary project, University of St.Gallen. https://sensing-group.com/files/theses/imp-haller-laenzlinger.pdf
[4] Länzlinger, J. (2025). Audio-centered Approach for Building a Multimodal Predictive AI Agent to Detect Depressive Behaviors. Master’s thesis, University of St.Gallen. https://sensing-group.com/files/theses/ma-jonas-laenzlinger.pdf
[5] Dupuis, K., & Pichora-Fuller, M. K. (2010). Toronto Emotional Speech Set (TESS). University of Toronto, Psychology Department.
[6] Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155-163.
[7] Bland, J. M., & Altman, D. G. (1986). Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet, 327(8476), 307-310.
Interested? Email the advisors: Bruno Rodrigues.