Diffusion Model
- In Turkish
- Difüzyon Modeli
In short
A diffusion model is a generative AI model that makes images, audio, or video by starting from random noise and removing it step by step until content appears.
What is a diffusion model?
A diffusion model is a type of generative AI model best known for creating images from text descriptions, and it is also used for video, audio, and 3D content. Instead of drawing a picture from left to right, it starts with an image of pure random noise, like television static, and gradually removes the noise over many steps until a detailed picture emerges.
Training works in reverse. The model is shown real images with increasing amounts of noise added, and a neural network learns to predict what noise was added at each level. At generation time, it repeatedly predicts and subtracts noise from a random starting point, guided by an embedding of the text prompt so the result matches the description. Many modern systems run this process in a compressed latent space rather than on full-size pixels, which makes it much faster, and the number of denoising steps trades quality against speed.
A good analogy is a sculptor who sees a figure hidden in a block of marble and chips away a little at a time until it appears. Diffusion models power text-to-image tools, image editing features such as filling in or extending parts of a photo, video generation, upscaling of low-resolution images, and even research tasks such as designing molecules.
Diffusion models are often confused with large language models. An LLM generates text one token at a time, each based on the ones before, while a diffusion model refines the whole output at once over many rounds. They are also different from GANs (generative adversarial networks), an older approach where one network generates images and another judges them; diffusion models are usually more stable to train and produce more varied results, but generation takes more steps.
Key takeaways
- A diffusion model generates content by removing noise step by step.
- It is trained by adding noise to real data and learning to predict that noise.
- A text prompt guides the denoising so the output matches the description.
- Many systems work in a compressed latent space for speed.
- Unlike an LLM, it refines the whole output at once rather than token by token.
Example
# Simplified generation loop of a text-to-image diffusion model.
# denoiser, text_encoder, decoder, and random_noise are placeholders for trained parts.
def generate(prompt, steps=30):
guidance = text_encoder(prompt) # embedding of the text prompt
image = random_noise(shape=(64, 64, 4)) # start from pure static
for t in reversed(range(steps)):
# Predict the noise still in the image, guided by the prompt
predicted_noise = denoiser(image, t, guidance)
image = image - predicted_noise / steps # remove a little of it
return decoder(image) # turn the latent back into a full-size pictureReaders ask
How do diffusion models generate images from text?
The text prompt is converted into an embedding that steers each denoising step. Starting from random noise, the model repeatedly removes noise in a way that makes the image more consistent with that embedding, until a finished picture remains.
What is the difference between a diffusion model and a GAN?
A GAN produces an image in one pass using a generator trained against a judging network, while a diffusion model builds the image over many denoising steps. Diffusion models are generally easier to train and produce more diverse images, while GANs generate faster.
Why is it called a diffusion model?
The name comes from diffusion in physics, where particles gradually spread out, like a drop of ink in water. Training gradually diffuses an image into noise, and the model learns to run that process backward.
See also
- Generative AIAI & Machine Learning, p. 20Generative AI is artificial intelligence that creates new content, such as text, images, code, or audio, based on patterns learned from existing data.
- Neural NetworkAI & Machine Learning, p. 33A neural network is a machine learning model made of layers of connected artificial neurons that learn patterns from data by adjusting numeric weights.
- Deep LearningAI & Machine Learning, p. 14Deep learning is a subset of machine learning that uses neural networks with many layers to learn complex patterns from raw data such as images and text.
- Multimodal AIAI & Machine Learning, p. 31Multimodal AI is artificial intelligence that can understand or generate several types of data, such as text, images, audio, and video, in a single model.
- Computer VisionAI & Machine Learning, p. 10Computer vision is the field of AI that enables computers to interpret images and video, such as recognizing objects, reading text, or detecting faces.
- EmbeddingAI & Machine Learning, p. 16An embedding is a list of numbers, called a vector, that represents the meaning of text, images, or other data so that similar items end up close together.
Spotted a mistake or something missing on this page?Suggest an edit