PHALAR: Phasors for Learned Musical Audio Representations
Learn how PHALAR, a novel framework, improves stem retrieval in musical audio representations by leveraging phasors and contrastive learning, achieving 70% higher accuracy with fewer parameters and faster training
- Implement a Learned Spectral Pooling layer to extract relevant spectral features from audio data
- Design a complex-valued head to capture temporal information in audio signals
- Apply contrastive learning to train PHALAR, leveraging phasors for improved representation learning
- Evaluate PHALAR's performance on stem retrieval tasks, comparing accuracy and training speed to state-of-the-art models
- Fine-tune PHALAR's parameters to optimize its performance on specific musical audio datasets
Audio engineers, music information retrieval researchers, and machine learning practitioners can benefit from PHALAR's innovative approach to stem retrieval, enabling more efficient and accurate music processing
💡 PHALAR's use of phasors and contrastive learning enables efficient and accurate stem retrieval, outperforming state-of-the-art models while requiring fewer parameters and less training time
Introducing PHALAR: a novel framework for musical audio representations, achieving 70% higher accuracy in stem retrieval with fewer parameters and 7x faster training! #musicinformationretrieval #machinelearning
Key Takeaways
Learn how PHALAR, a novel framework, improves stem retrieval in musical audio representations by leveraging phasors and contrastive learning, achieving 70% higher accuracy with fewer parameters and faster training
Full Article
Abstract:
arXiv:2605.03929v2 Announce Type: cross Abstract: Stem retrieval, the task of matching missing stems to a given audio submix, is a key challenge currently limited by models that discard temporal information. We introduce PHALAR, a contrastive framework achieving a relative accuracy increase of up to $\approx 70\%$ over the state-of-the-art while requiring $<50\%$ of the parameters and a 7$\times$ training speedup. By utilizing a Learned Spectral Pooling layer and a complex-valued head, PHALAR en
DeepCamp AI