← Theses

An Array That Knows When It Is Wrong

Open topic · BA · Advisors: Bruno Rodrigues

Motivation

A microphone reports that someone’s pitch was 142 Hz. Was it? In a laboratory you would know, because you chose the recording and can check. In a home you cannot, because nobody knows what the true pitch of a passing sentence was and no reference exists to check against. This is the awkward fact underneath in-home vocal sensing, since the system produces numbers all day and has no way to separate the good ones from the bad. A measurement taken while a tap was running, or from four metres away through a doorway, arrives in the database looking exactly like a clean one.

Problem statement

Our array suggests a way out, because it has eight nodes and they overhear one another. When several capture the same utterance, each extracts its own estimate of pitch, jitter, and shimmer, and if those estimates agree the measurement is probably sound. This is the useful asymmetry: disagreement can be observed at runtime, in any house, with no reference, while error cannot. Whether the first predicts the second has not been established for vocal features, and it is not obvious that it holds, since nodes can agree and be wrong together. If it does hold, the array can attach a confidence score to every value it reports, computed from nothing but its own internal consistency, which is the missing piece for any clinical use of this data.

The project

We will test it. Replay supplies the truth a home does not, since a known recording played into the room has feature values computable from the original file, so every co-captured utterance yields both the disagreement among nodes and the true error of each, and fitting one against the other gives a calibration curve. The nodes, the server, and the feature extraction already run, and the grouping of simultaneous captures is implemented, so the work starts at the experiment rather than at the plumbing. Reference speech comes from the TESS corpus, so the study needs no participants and touches no personal data. No background in speech processing or statistics is required, and the hardware is provided. The specific goal of this project is to determine whether disagreement between co-located nodes predicts the error of a vocal feature measurement, and to turn that relationship into a calibrated confidence score an in-home array can compute for itself without ground truth.

Objectives

  • Understand the IHearYou signal chain and why in-home vocal measurement has no reference.
  • Group simultaneous captures across co-located nodes and extract vocal features independently per device.
  • Quantify inter-device agreement per feature using intraclass correlation and Bland-Altman limits of agreement.
  • Obtain absolute error by replaying known reference clips through a speaker in the room and scoring each node against the original file.
  • Test whether disagreement predicts error, build a calibrated confidence score, and report its calibration error and its failure cases.

Previous theses in this line

References

[1] J. Länzlinger, K. O. E. Müller, B. Stiller, and B. Rodrigues. Towards Interpretable Depression Detection: Linking Acoustic Features to DSM-5 Indicators. IEEE PerCom 2026, Work-in-Progress.

[2] Khamaisi, K., Keller, N., Krummenacher, S., Huber, V., Fässler, B., & Rodrigues, B. (2025). From Noise to Knowledge: Acoustic Anomaly Detection in Pumped-storage Hydropower Plants. arXiv:2509.22881.

[3] Länzlinger, J. (2025). Audio-centered Approach for Building a Multimodal Predictive AI Agent to Detect Depressive Behaviors. Master’s thesis, University of St.Gallen. https://sensing-group.com/files/theses/ma-jonas-laenzlinger.pdf

[4] Haller, T. (2026). Explainable Multimodal-based Depression Awareness at the Edge. Master’s thesis, University of St.Gallen. https://sensing-group.com/files/theses/ma-tibor-haller.pdf

[5] Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155-163.

[6] Bland, J. M., & Altman, D. G. (1986). Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet, 327(8476), 307-310.

[7] Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. Proc. ICML 2017, 1321-1330.

[8] Dupuis, K., & Pichora-Fuller, M. K. (2010). Toronto Emotional Speech Set (TESS). University of Toronto, Psychology Department.

Interested? Email the advisors: Bruno Rodrigues.