Computer Vision
- In Turkish
- Bilgisayarlı Görü
In short
Computer vision is the field of AI that enables computers to interpret images and video, such as recognizing objects, reading text, or detecting faces.
What is computer vision?
Computer vision is the branch of artificial intelligence that teaches machines to understand visual input, meaning photos, video, and live camera feeds. Typical tasks include image classification (what is in this picture?), object detection (where is each object?), segmentation (which exact pixels belong to each object?), and optical character recognition, or OCR, which reads text from images.
To a computer, an image is just a grid of numbers, with one value per pixel for brightness or three values for red, green, and blue. Modern computer vision uses deep learning: convolutional neural networks (CNNs) scan small patches of the image to detect edges, then shapes, then whole objects, and vision transformers treat patches of an image like tokens in a sentence. These models learn from large sets of labeled images rather than from hand-written rules.
Think of how a child learns to recognize a cat: nobody gives them a formula, they simply see many cats until the pattern clicks. Computer vision powers phone camera features, face unlock, barcode and document scanners, medical image analysis, quality checks in factories, and the cameras in driver-assistance systems.
Computer vision is sometimes confused with image processing. Image processing transforms an image, for example by resizing, sharpening, or applying a filter, while computer vision extracts meaning from it, such as 'this photo contains two dogs'. Many vision systems use image processing as a first step before a model interprets the result, and multimodal models that accept both text and images now blur the line between computer vision and natural language processing.
Key takeaways
- Computer vision lets software extract meaning from images and video.
- Core tasks are classification, object detection, segmentation, and OCR.
- Images are grids of pixel values processed by neural networks such as CNNs and vision transformers.
- Image processing edits pixels, while computer vision interprets them.
Example
# A tiny 3x3 grayscale image is just numbers (0 = black, 255 = white)
image = [
[0, 0, 0],
[255, 255, 255],
[0, 0, 0],
]
# A hand-written "feature": does the image contain a bright horizontal line?
def has_horizontal_line(img):
return any(all(pixel > 200 for pixel in row) for row in img)
print(has_horizontal_line(image)) # True
# Neural networks learn thousands of features like this from labeled examplesReaders ask
What is the difference between computer vision and image processing?
Image processing changes an image, such as resizing, cropping, or adjusting brightness, while computer vision interprets it, such as identifying objects or reading text. Vision systems often use image processing to prepare images before analyzing them.
What is object detection?
Object detection is a computer vision task that finds each object in an image and draws a labeled bounding box around it, such as 'car' or 'person'. It goes further than classification, which only says what the image contains overall.
Is computer vision part of machine learning?
Modern computer vision is built almost entirely on machine learning, especially deep learning. Older approaches used hand-designed rules and filters, which still appear in simple or resource-limited systems.
See also
- Deep LearningAI & Machine Learning, p. 14Deep learning is a subset of machine learning that uses neural networks with many layers to learn complex patterns from raw data such as images and text.
- Neural NetworkAI & Machine Learning, p. 33A neural network is a machine learning model made of layers of connected artificial neurons that learn patterns from data by adjusting numeric weights.
- Machine LearningAI & Machine Learning, p. 27Machine learning is a branch of artificial intelligence in which computers learn patterns from data to make predictions instead of following hand-written rules.
- TransformerAI & Machine Learning, p. 49A transformer is a neural network architecture that uses attention to weigh how each token in a sequence relates to the others, and it powers most modern LLMs.
- Natural Language ProcessingAI & Machine Learning, p. 32Natural language processing is the field of AI that teaches computers to read, understand, and generate human language in the form of text or speech.
- InferenceAI & Machine Learning, p. 24Inference is the stage where a trained machine learning model is used to make predictions or generate output from new data, without changing what it learned.
Spotted a mistake or something missing on this page?Suggest an edit