Nvidia released a paper about a 100KB text-to-image model that only trained for 4 minutes but claims to be better than bigger models

hayek@feddit.de · 1 year ago

Nvidia released a paper about a 100KB text-to-image model that only trained for 4 minutes but claims to be better than bigger models

ubermeisters@lemmy.world · 1 year ago

Pretty neat. The training process takes a while for textual inversion, which I have enjoyed playing around with. I hope Automatic1111 gets support for this method of training, if it takes off!

AngrilyEatingMuffins@kbin.social · 1 year ago

Can this be adapted to LLMs?

ubermeisters@lemmy.world · 1 year ago

Great question, I wondered the same thing. I’ve got a decent knowledge base where stable diffusion (text to image etc) is concerned, and understand the applications of this Nvidia process, I’m not familiar enough with customization options for LLMs. I haven’t really seen references to hypernetwork/lora/midjourney type applications in LLMs, or anything that really “plugs into” your existing model to augment results, the way stable diffusion is geared for customization. It seems in my limited understanding, that customization for LLMs requires customization of the training ing data, and a completely new training process for the actual model, not a reference model like SD.

Nvidia released a paper about a 100KB text-to-image model that only trained for 4 minutes but claims to be better than bigger models

Nvidia released a paper about a 100KB text-to-image model that only trained for 4 minutes but claims to be better than bigger models

Key-Locked Rank One Editing for Text-to-Image Personalization