I've been having a hard time finding a solution. Is there really no commercialized model that I can feed in a reference sound and text instruction and get another sound out? Right now having a multimodal inputs to image or text output is a commodotized, solved problem.
Source: [Hacker News](https://news.ycombinator.com/item?id=49972125)