Independent research
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
SwanTale is a unified model for multi-speaker speech and audio generation supporting both zero-shot and instruct tasks. It introduces SwanData-Caption, a data pipeline that cleans raw audio, adds targeted synthetic coverage (elderly speech, short utterances, challenging pronunciations), and annotates multi-level captions (environment, speakers, content).…
Yu Zhang, Ruiqi Li, Changhao Pan, Ke Lei, et al.- Published
- Aug 2026
- Citations
- 0
- Code
- Not linked
