Transcript
[00:00] So that is fucking awful. I’m sorry I made that. So this is a step 1.5, which is really cool. It’s got a whole UI and it’s awesome. And you should check it out. Now this is a step C++, which is a fork. I’m sorry, not a fork. It’s a C++ implementation of all the stuff that’s in a step 1.5. And it’s really cool, but it’s super crude. And then this is hot step, which is what we’re going to talk about, which is a fork of a step C++.
[00:48] So to get started, you download the latest release, which is going to be some form of compressed zip file or whatever. There’s instruction for Windows and Linux and Mac. I’m on a Mac. So I download it, I extract it, and then I run this shell command. So we’re going to do that here. I’m in the directory where I have all that. I’m going to run this guy. And as it starts up, I just want to point out that it actually gives you two URLs. One of them is the a server URL. Just open that for a second here. And this is the ace step C++ UI. It’s pretty crude, but it’s a lot of fun. But we’re going to focus on hot steps. So let’s open that guy up. Here is our UI.
[01:36] So once we’re in here, the first thing we get is this Autogen UI. Now I’m going to jump up here because as soon as you load this, what you’re wanting to do is download a bunch of models. So I’ve downloaded a whole bunch. I don’t really know what I’m doing. But I just grabbed a bunch of stuff. But once you have those, right off the bat, you can jump into this Autogen where you can create an instrumental or just lyrics or lyrics and AI. So I’m going to come in here. We’ve got a whole bunch of genres. I’m going to say at random, see if something cool comes up. Doom metal, blues, pop punk, modern jazz. You know what? Let’s go for it. We’re going to give it a subject. I’m going to say coding makes my hands hurt. Now I’ve got this connected to a llama, which you can update in your settings. And once you’re connected, you can pick any model. I’m just using llama three, eight billion parameters. It’s not always the best, but we’re just going to see how it goes.
[02:29] So what I’m going to do here is I’m going to generate lyrics and that is going to use the large language model to generate not only the lyrics, but again, the tone, the style, the time signature, beats per minute. And obviously, yes, the lyrics. Okay. So llama is not that great at this. If we jump up to, oh, we don’t have it here, but there’s a system prompt that says like only return certain things. And where it can screw up is it starts putting all of this description stuff in the actual lyrics and that’s no good. So I’m going to jump over to refine and custom gen. I’m going to take this section in tags and drop it in there. And then I’m just going to clean all this up. I don’t think I care about the commas at the end. There’s a chorus, instrumental parts, breakdown, and we’ve got a title. So let’s grab that. Drop that into the title. Now we have beats per minute and key. So beats per minute is 120. The duration is going to be 270. The key is C minor. And the time signature is 44. I think that’s it. So I can get rid of this stuff. Cool. So you don’t always have to do that. And with a better model, it’s definitely not going to happen that way. Usually it’ll just populate it for you. But if you do run into a situation where you get just all of it back in the lyrics, that’s all you need to do is jump over refine and parse all that out yourself. So we’re going to hit generate and see what we get.
[04:23] All right, that took about three minutes. We are going to press play and see what we get. All right, I started it too soon. Let’s go ahead and give it a second. Coding all night, coding all day. My hands ache, my mind astray. I’m lost in a sea of code. Let’s jump into the middle somewhere and just see how this sounds. Let’s back up and hear how that goes. All right. I don’t think it’s a hit, but it’s pretty cool. Okay, so that is how to do auto gen and custom gen.
[05:41] We’re going to jump into this lyric studio part. Now, I’ve been messing with some stuff. Black Sabbath, Gravediggers, you know, and it’s got all the lyrics. So the way this works is if I want to go back and say, add an artist and I say, fetch from genius, I’m going to try to fetch toadies. Cool. So we now have toadies and I can go into this album and I can see the lyrics to all the songs. So you’ll notice, let’s see. So there’s bridge sections, chorus sections. There’s often sections for instrumental stuff, guitar solos and things like that. So it’s super helpful if you want to use this for something.
[06:29] So another really cool feature is artist profiles. Now, obviously, I don’t have one for toadies, but I do have one for Sabbath. These take a really long time to generate, just so you know, or at least it did for me. But what it does is it breaks down the entire album into themes, subjects, you know, different categories, the tone, the vocabulary that the lyrics use, structural patterns for the song, all sorts of stuff. And then you can use that to further refine it and create new tracks based on it.
[07:05] Now, one that I’ve really enjoyed is this stem separator. Now, you do need to have a base model, so it’s not none of the turbo models will work here. But the way this works is you load up an audio file. Then you pick which elements from the audio file you would like to extract. And then, again, it takes a long time, but you end up with a whole bunch of tracks that you can potentially repurpose or use or train your models further on. So I did one on toadies. And one thing that helps is if you take those lyrics that you got from the Lyric Studio and you drop them right here. And those are the lyrics for the track that you add. Then you select what you want. Now, I did everything. I selected everything for this toadies track. And I don’t know if that’s the right thing or not because I’m just messing around. Now, one thing I’m going to do here is I’m going to turn all these down to 10% because I’m trying to capture this audio at the same time as my mic and all of that. And it’ll end up being really, really noisy. Okay, all those are turned down. And what I’m going to do, I’m actually going to turn the vocals up a little. I’m going to try 20. Hopefully, it’s not too loud. I’m going to hit preview on this and start playing. I’m going to jump to, I think I know a good part of the song to jump to. And hopefully, I don’t know how YouTube works with stuff like this. I hope I don’t get in trouble for playing a song.
[09:10] Okay, so that’s a really loud part of the song. It was right around here. I’m going to do this one more time. Do you want to be you want to be my angel? Give it up to me Give it up to me Do you want to be my angel? Give it up to me So that’s been really fun. I did one with War Pigs. I’m going to see how loud this sounds. I’ve only got two tracks on the drums. They love the drums. And the vocals. Death and hatred to mankind. Poisoning their brainwashed minds. Still sounds great. So honestly, it’s kind of metallic. There’s a little bit of messiness in there. But you could certainly take that and refine it further. Run it through. Denoisers, background cleanup, all of that. There’s a bunch of post-processing that I have turned off right now. When you turn that on, it is going to take longer to do some of these things. But you will get much better results.
[10:35] There’s also a giant mess of knobs and dials for various generation. Attributes. I haven’t dug into what all those are. But look, there’s a denoiser. So maybe that would have made those tracks sound better. You got your stem separator. Then you got your stem builder. Which you could then take those individual tracks and start to create new tracks from them. So you could load up the vocals from one song. Say, I’m targeting the vocals and start messing with that. There’s repaint that I have not jumped into. And it does say the component is under heavy development. So best of luck to you on that. There’s a groovy cover studio, which is, I’m not going to do it right now. But I was messing with taking the Andrews sisters, Straighten Up and Fly Right, and having the Gravediggers sing it, which was interesting. I don’t think it turned out great. But there’s a lot to mess around with here, where once you’ve built that profile of your artist, and then you take a song that you want to cover, presumably what you should be able to do is say, have Black Sabbath do a cover of this Andrews sister’s song. And it’s a lot of fun to mess with.
A feature-rich web application for local AI music generation powered by GGML, with native safetensors support. Describe a song with a text caption and lyrics, and get stereo 48kHz audio generated entirely on your local hardware — no cloud, no API keys, no subscriptions. It extends the acestep.cpp inference engine with 100+ features across inference, audio processing, and creative tooling.
Details
| URL | https://github.com/scragnog/HOT-Step-CPP |
| Type | Web App / CLI |
| Pricing | Free (open source) |
| Open Source | Yes |
| License | MIT (engine/), other licenses for plugins and models |
| Tech Stack | C++ (engine), CUDA/Vulkan/Metal (GPU acceleration), GGML, Node.js / TypeScript (server), React / Vite / TypeScript (UI), Lua (plugins), Essentia (audio analysis) |
| Platforms | Windows (x64: CUDA, Vulkan, CPU), Linux (x64: CUDA, Vulkan, CPU), macOS Apple Silicon (M1/M2/M3/M4: Metal) |
| Self-Hosted | Yes |
Key Features
- Full music generation suite — AceStep 1.5 and MiniMax-Music3 backends (first non-Python MM3 implementation), 17 solvers, 9 schedulers, 7 guidance modes, LoRA adapters with runtime mode, 100+ GGUF models across 5 HuggingFace repos with in-app Model Manager, and split-model format for MiniMax-Music3 (mix precision per pipeline component)
- Creative studios — Auto-Gen (AI-driven full song creation with genres, lyrics, metadata), Custom-Gen (full manual control), Cover Studio (reference track analysis + style matching + pitch/tempo), Stem Studio (4-stage neural separation: BS-RoFormer, Mel-Band RoFormer, MDX23C, HTDemucs), Stem Builder (generative new instrument stems), MIDI Studio (audio-to-MIDI via MuScriptor port), Repaint Studio (region-based regeneration), and StableStep (post-processing refiner using Stable Audio 3)
- Audio processing & plugins — Matchering mastering engine (loudness/EQ/dynamics matching to reference at 48kHz), VST3 host (scan/load/run plugins in generation pipeline), 100+ extensible Lua plugins (ODE/SDE solvers, schedulers, guidance modes, postprocess pipelines), spectral denoiser, PP-VAE/ScragVAE neural polish, vocal naturalizer, audio quality evaluator, and lossless WAV32 pipeline with WAV/MP3/FLAC export. Includes experimental Training Studio for training style adapters entirely within the app (LM LoRA + DiT LoRA)
Best For
Music producers, hobbyists, and researchers who want to generate music entirely locally using AI models without cloud services or API keys. Particularly useful for users with NVIDIA/AMD/Intel GPUs (Windows/Linux) or Apple Silicon Macs who want a rich feature set covering generation, remixing, stem separation, mastering, and audio-to-MIDI transcription — all on their own hardware.
Integrations
acestep.cpp (inference engine), GGML (model format), MiniMax-Music3 (backend), Essentia (audio analysis), VST3 plugins, MuScriptor (audio-to-MIDI), SuperSep (stem separation: BS-RoFormer, Mel-Band RoFormer, MDX23C, HTDemucs), Stable Audio 3 (StableStep refiner), HuggingFace (model repository), Lua (plugin system), 7 LLM providers (Gemini, LM Studio, OpenAI-compatible) for Lyric Studio.
Notes
- Pre-built portable releases available — no installation required, just extract and run
- Platform requirements: Windows 10/11 64-bit or Linux Ubuntu 22.04+ or macOS 13+ Apple Silicon. ~10 GB disk space. GPU variants need CUDA 12.x+ (NVIDIA) or Vulkan 1.1+ (AMD/Intel/NVIDIA). macOS requires Xcode (not just CLI tools)
- Node.js 18-22 LTS required (Node 24+ not supported). Build from source requires Visual Studio 2022 (Windows), CMake 3.14+, Git
- Recommended model sizes: LM 4.2 GB, Text Encoder 748 MB, DiT 2.4 GB, VAE 322 MB (ACE-Step 1.5). MiniMax-Music3: ~13 GB (Q8_0) or ~24 GB (F16)
- Training Studio is highly experimental — GPU-hungry (16 GB+ VRAM recommended, 24 GB+ for full DiT training)
- Community: Discord server (primary support channel), Buy Me a Coffee (donations), HuggingFace (models)
- macOS releases are unsigned — may need xattr -cr to remove quarantine flag on first launch