Interpretable depression detection from speech: work-in-progress at IEEE PerCom 2026
Our work-in-progress paper Towards Interpretable Depression Detection: Linking Acoustic Features to DSM-5 Indicators was presented at IEEE PerCom 2026, the 24th IEEE International Conference on Pervasive Computing and Communications, held from 16 to 20 March 2026 in Pisa, Italy. PerCom is the premier venue in pervasive and ubiquitous computing and one of the few conferences rated A* (CORE), so being accepted into its Work-in-Progress track is an encouraging early signal for the direction. The work is a collaboration with the Communication Systems Group (CSG) at the University of Zurich, and was presented by Bruno Rodrigues.

Motivation. Depression affects more than 280 million people worldwide, yet diagnosis still leans on subjective self-reports and episodic clinical interviews that can miss how a person actually behaves day to day. Machine learning can detect depression from speech with high accuracy, but the models usually return a single opaque score: a clinician cannot see which symptoms drove it, and streaming raw audio to the cloud is a serious privacy risk inside the home.
The idea. The paper introduces a transparent Linkage Framework that maps interpretable acoustic features (pitch variability, pause duration and frequency, speech tempo, jitter and shimmer) directly onto DSM-5 depressive-behavior indicators through explicit, testable rules instead of a black box. Each feature-to-indicator link is grounded in the clinical literature, so the output is a symptom-level readout a clinician can inspect rather than just a probability. The whole pipeline is designed to run locally on commodity hardware, so raw speech never leaves the household.
Progress so far. A preliminary evaluation on the DAIC-WOZ corpus (64 participants) found directionally consistent associations for reduced pitch variability, longer pauses, and slower speech tempo, matching the hypothesized indicators for psychomotor change and concentration difficulty. On a MacBook Pro M1, the pipeline sustained real-time throughput with latency below one second per ten-second window and without a GPU, confirming that interpretable, on-device analysis is feasible. Effect sizes remain moderate on this small, single-session dataset, which is precisely why the work is framed as work-in-progress: the next steps are longitudinal validation with naturalistic home speech, multimodal fusion with wearable signals, and calibration across languages and demographics.
Roots and next steps at HSG. The framework grew out of the master thesis of Jonas Länzlinger at HSG, Audio-centered Approach for Building a Multimodal Predictive AI Agent to Detect Depressive Behaviors, and it complements the group’s network-side work on the same clinical target, CareNet, presented at NOMS 2026. The line continues in an open master thesis that asks how language-independent these privacy-preserving speech features really are.
Copy & share
Edit if you like, then copy and paste into your post.