🎙️ AuK — speech generation & editing

One 1.5B model, one instruction box. AuK does zero-shot & instruct TTS, content / lyric editing, pitch–speed–volume editing, emotion / timbre / accent / nonverbal / whisper edits, denoising and source separation — all from natural language, in English or Chinese.

Drop in a clip (or leave it empty for instruct TTS), describe the edit, set how long the result should be, and hit Generate.

Model
0 30

Tips · Target duration is what AuK renders: 0 copies the source clip's length — right for enhancement, separation, emotion or accent edits. Set it explicitly when you change what is said (TTS, content editing) or how fast (speed editing) — roughly the length you expect the new audio to take. Flash samples in 4 steps and is ~6× faster; Base trades speed for fidelity.

Examples

Examples

Model: tencent/AuK · tencent/AuK-Flash · text encoder Qwen/Qwen2.5-Omni-3B · ComfyUI-packaged weights: drbaph/AuK-comfyui · code Tencent-Hunyuan/AuK (MIT). Example audio is taken unchanged from the AuK repository.