Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment
The study introduces a method to ground respiratory sound encoders in clinical semantics by leveraging an LLM to produce structured medical reports from available metadata, forming semantically rich anchors for training.
Its training blends a sigmoid-based contrastive loss with the encoder’s self-supervised objective and similarity-aware negative sampling, enabling sharper discrimination of pathological vs. non-pathological states.
Evaluated across nine tasks on six datasets, the method reports a 61.3% mean zero-shot AUC, outperforming prior models and achieving the highest linear probing AUC (71.6%) using only 43% of the data used by larger baselines.
Why it matters: Demonstrates that structured semantic alignment via LLM-generated anchors can improve zero-shot clinical diagnostics, potentially reducing labeled data needs for medical audio tasks.
AI-assisted brief
official source
Add a comment
No account required