Exploring Token-Space Manipulation in Latent Audio Tokenizers

📰 ArXiv cs.AI

arXiv:2605.11192v1 Announce Type: cross Abstract: Neural audio codecs provide compact discrete representations for speech generation and manipulation. However, most codecs organize tokens as frame-level sequences, making it difficult to study or intervene on global factors of variation. In this work, we propose the Latent Audio Tokenizer for Token-space Editing (LATTE) that appends a fixed set of learnable latent tokens to the audio feature sequence and retains only these tokens for quantization

Published 13 May 2026
Read full paper → ← Back to Reads