Tencent’s AuK handles 16 audio tasks from one prompt
Tencent’s Hunyuan team, working with Shanghai Jiao Tong University and the Shanghai Innovation Institute, has released AuK on GitHub. The open-source audio foundation model uses a shared natural-language interface for voice cloning, speech generation, content editing, denoising, source separation, acoustic controls, and voice transformations. Its core generator has 1.5 billion parameters and depends on a separately downloaded Qwen2.5-Omni-3B model for semantic conditioning.
Conventional speech pipelines often combine text-to-speech, voice conversion, denoising, editing, and separation services. AuK moves those operations behind one message format, reducing the routing and integration code required to connect specialist systems. Developers provide an instruction, attach source or reference audio when needed, and receive a generated waveform.
Sixteen jobs, one message format
The AuK project site describes a chat-style API containing text and optional audio messages. A source clip provides material to transform, while a reference clip supplies characteristics such as speaker identity or style.