Skip to main content

Computer Vision

Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/computer-vision

In short

Computer vision is the field of AI that enables computers to interpret images and video, such as recognizing objects, reading text, or detecting faces.

What is computer vision?

Computer vision is the branch of artificial intelligence that teaches machines to understand visual input, meaning photos, video, and live camera feeds. Typical tasks include image classification (what is in this picture?), object detection (where is each object?), segmentation (which exact pixels belong to each object?), and optical character recognition, or OCR, which reads text from images.

To a computer, an image is just a grid of numbers, with one value per pixel for brightness or three values for red, green, and blue. Modern computer vision uses deep learning: convolutional neural networks (CNNs) scan small patches of the image to detect edges, then shapes, then whole objects, and vision transformers treat patches of an image like tokens in a sentence. These models learn from large sets of labeled images rather than from hand-written rules.

Think of how a child learns to recognize a cat: nobody gives them a formula, they simply see many cats until the pattern clicks. Computer vision powers phone camera features, face unlock, barcode and document scanners, medical image analysis, quality checks in factories, and the cameras in driver-assistance systems.

Computer vision is sometimes confused with image processing. Image processing transforms an image, for example by resizing, sharpening, or applying a filter, while computer vision extracts meaning from it, such as 'this photo contains two dogs'. Many vision systems use image processing as a first step before a model interprets the result, and multimodal models that accept both text and images now blur the line between computer vision and natural language processing.

Key takeaways

  • Computer vision lets software extract meaning from images and video.
  • Core tasks are classification, object detection, segmentation, and OCR.
  • Images are grids of pixel values processed by neural networks such as CNNs and vision transformers.
  • Image processing edits pixels, while computer vision interprets them.

Example

How a computer sees an imagepython
# A tiny 3x3 grayscale image is just numbers (0 = black, 255 = white)
image = [
    [0, 0, 0],
    [255, 255, 255],
    [0, 0, 0],
]

# A hand-written "feature": does the image contain a bright horizontal line?
def has_horizontal_line(img):
    return any(all(pixel > 200 for pixel in row) for row in img)

print(has_horizontal_line(image))  # True
# Neural networks learn thousands of features like this from labeled examples

Readers ask

What is the difference between computer vision and image processing?

Image processing changes an image, such as resizing, cropping, or adjusting brightness, while computer vision interprets it, such as identifying objects or reading text. Vision systems often use image processing to prepare images before analyzing them.

What is object detection?

Object detection is a computer vision task that finds each object in an image and draws a labeled bounding box around it, such as 'car' or 'person'. It goes further than classification, which only says what the image contains overall.

Is computer vision part of machine learning?

Modern computer vision is built almost entirely on machine learning, especially deep learning. Older approaches used hand-designed rules and filters, which still appear in simple or resource-limited systems.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings