Multimodal AI
- In Turkish
- Çok Modlu Yapay Zekâ
In short
Multimodal AI is artificial intelligence that can understand or generate several types of data, such as text, images, audio, and video, in a single model.
What is multimodal AI?
A modality is a kind of data, such as text, images, audio, or video. Early AI systems usually handled only one: a language model read text, and a computer vision model looked at pictures. Multimodal AI combines several modalities, so a single model can, for example, answer a question about a photo, describe a chart, summarize a meeting recording, or generate an image from a written description.
Most multimodal models convert each type of input into embeddings, lists of numbers that capture meaning, in a shared space the model can reason over. An image is split into small patches and a sound clip into short frames, and each piece becomes a token much like a word in a sentence. A transformer then processes all these tokens together, which lets the model connect the word dog in a question with the dog in a picture.
Think of the difference between reading a transcript of a video call and actually watching it: seeing faces and slides and hearing tone of voice gives far more context. Multimodal AI is used in document processing that reads scanned forms and tables, accessibility tools that describe images for blind users, voice assistants, visual quality inspection in factories, and coding assistants that turn a screenshot of a design into code.
Multimodal AI is often confused with a pipeline of separate single-purpose models, such as speech-to-text feeding a text-only chatbot and then a text-to-speech engine. That chaining works, but information like tone of voice or image layout is lost between steps, while a natively multimodal model processes the modalities together. Multimodal input and multimodal output are also different: many models accept images but can only answer in text.
Key takeaways
- A modality is a type of data, such as text, images, audio, or video.
- Multimodal models map different data types into a shared embedding space.
- They can answer questions about images, read documents, and handle speech.
- Accepting several input types doesn't mean a model can produce them all.
- Natively multimodal models keep context that chained single-mode models lose.
Example
// Ask a multimodal model about an image (endpoint and fields are illustrative)
const response = await fetch("https://api.example.com/v1/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
messages: [{
role: "user",
content: [
{ type: "image", url: "https://example.com/receipt.jpg" },
{ type: "text", text: "What is the total amount on this receipt?" },
],
}],
}),
});
console.log((await response.json()).text);Readers ask
What is the difference between multimodal AI and computer vision?
Computer vision focuses on understanding images and video, often with models built for one task like detecting objects. Multimodal AI combines vision with other modalities, such as language, so a single model can look at an image and discuss it in natural language.
Are large language models multimodal?
Many modern ones are. Text-only LLMs process only text, but a growing number of models also accept images, audio, or video as input, and some can generate images or speech as output.
What is a vision-language model?
A vision-language model, or VLM, is a multimodal model that takes images and text as input and usually produces text. It is used for tasks like image captioning, visual question answering, and reading documents.
See also
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- Computer VisionAI & Machine Learning, p. 10Computer vision is the field of AI that enables computers to interpret images and video, such as recognizing objects, reading text, or detecting faces.
- EmbeddingAI & Machine Learning, p. 16An embedding is a list of numbers, called a vector, that represents the meaning of text, images, or other data so that similar items end up close together.
- TransformerAI & Machine Learning, p. 49A transformer is a neural network architecture that uses attention to weigh how each token in a sequence relates to the others, and it powers most modern LLMs.
- Generative AIAI & Machine Learning, p. 20Generative AI is artificial intelligence that creates new content, such as text, images, code, or audio, based on patterns learned from existing data.
- Natural Language ProcessingAI & Machine Learning, p. 32Natural language processing is the field of AI that teaches computers to read, understand, and generate human language in the form of text or speech.
- Diffusion ModelAI & Machine Learning, p. 15A diffusion model is a generative AI model that makes images, audio, or video by starting from random noise and removing it step by step until content appears.
Spotted a mistake or something missing on this page?Suggest an edit