Quick answer

Computer vision is the field of artificial intelligence in which a system analyzes images or video and understands what they show: it recognizes objects, locates them, measures them and reads text. Modern systems learn this from thousands of labeled examples rather than from hand-written rules. The breakthrough came in 2012, when the AlexNet neural network reached a top-5 error of 15.3% in the ImageNet competition, compared with 26.2% for the best classic approach (Krizhevsky et al., 2012).

What exactly is computer vision?

It is technology that turns pixels into information you can act on. A person looks at a photo and instantly sees a car, a cracked tile or an invoice number. To a computer it is an array of numbers in which a model has to find the meaning.

Two terms are often confused:

  • Image processing changes pixels using fixed rules: it corrects perspective, improves contrast, detects edges.
  • Computer vision interprets content: what is in the image, where it is and in what condition.

In real systems both work together. Take document capture from phone photos: the system detects the edges of the page, corrects the perspective, reads the text and moves values into form fields, and a person only checks the result.

The business value of computer vision is not that a machine "sees better" than a person. It is that it watches all the time, to the same standard, on every item rather than a sample.

How does a machine learn to see?

From examples. Instead of writing rules that describe what a scratch on paint looks like, you show the model many photos with and without scratches, and it finds the features that tell them apart. This is called supervised learning: each example has a label, the correct answer.

The process looks like this:

  1. Collect and label data. Photos from real conditions, annotated by people: "defect / no defect", a box around an object or an outline pixel by pixel.
  2. Learn patterns. Layer by layer, a neural network detects increasingly complex features: first edges and colors, then textures and shapes, finally whole objects.
  3. Test on new data. The model is evaluated on images it did not see during training. Only this result tells you whether it will cope in production.
  4. Deploy and monitor. The model runs on a camera, a server or in an app, and its mistakes come back as new examples for the next training round.

Data scale mattered enormously. ImageNet, the dataset behind the breakthrough models, contains more than 14 million labeled images in more than 21,000 categories (ImageNet). In a company you rarely start from scratch: you take a model pre-trained on a large dataset and fine-tune it on your own images. That is what made computer vision accessible to mid-sized businesses.

Which architectures power modern systems?

For a decade convolutional neural networks (CNNs) dominated. They slide small filters across the image and are good at catching local patterns. Since 2020 vision transformers (ViT) have played a growing role: they split an image into patches and analyze relationships between them, much like language models analyze words (Dosovitskiy et al., 2020).

The newest direction is vision-language models (VLMs), which connect images with text: you can ask what is in a photo or have a chart described. We use this in our agentic knowledge base: charts in company documents are described by a vision model, so their content can later be searched and cited.

What tasks does computer vision solve?

Most business applications come down to four core tasks. They differ in what you ask of the image, how much data preparation costs and how quality is measured.

TaskQuestion it answersBusiness exampleHow data is labeled
ClassificationWhat is in this image?Good or defective product, document typeOne label per image
Object detectionWhat and where?Vehicles at a gate, a hard hat on a worker, products on a shelfA box around each object
SegmentationWhich exact pixels belong to the object?Outline of a scratch, damaged area, a finding in a medical imageA pixel-level outline
OCRWhat text is in the image?Invoices, forms, notes from a photo, license platesText assigned to an image region

Classification is the simplest and cheapest in terms of data. It is enough when you only need a yes/no answer or a category.

Detection adds location. It lets a system count objects, check whether someone entered a restricted zone, or read a license plate from a frame that also shows the street and sidewalk.

Segmentation is the most precise and the most expensive to label, because a person outlines every object. Foundation models help here, such as Segment Anything, trained on more than a billion masks from 11 million images, which can propose an object outline from a single click (Kirillov et al., 2023). People correct the proposal instead of drawing from scratch.

OCR (optical character recognition) is a separate and very practical branch: turning an image of text into text you can search and process. We cover how OCR evolved from simple tools to models that understand document layout in our article on modern OCR for business.

These tasks are chained together. An access control system first detects a vehicle, then crops the license plate, and finally reads it with OCR and checks it against an allowlist.

Where is computer vision already working?

Anywhere someone looks at something many times a day and judges it against repeatable criteria. The most common areas are manufacturing quality control, workplace safety, logistics and warehousing, retail, agriculture, documents and healthcare.

Healthcare shows how mature the technology is. The FDA's list of AI-enabled medical devices has more than 1,600 entries, about three quarters of them in radiology, which is image analysis (our count of the list as of September 4, 2026; the FDA notes that the list is not comprehensive) (FDA). We walk through concrete business applications, with numbers and an approach to ROI, in 5 ways computer vision solves business problems.

What are the limits of computer vision?

A model is only as good as the data it learned from, and only in conditions similar to the training ones. This is the single most important thing to know before your first project.

  • Domain shift. A new camera, different lighting, new packaging or a change of season can noticeably lower the accuracy of a model that looked great in testing. Test data must come from the real environment, and the model must be monitored after deployment.
  • Rare cases. Defects that occur rarely are seen rarely and recognized worse. They have to be collected deliberately, and sometimes generated synthetically.
  • Label quality. If two people disagree about what counts as a defect, the model will learn that inconsistency. Write down the criteria before labeling starts.
  • No common sense. A model does not "understand" the world. It can make mistakes a person never would and does not always signal that it is unsure.
  • Cost of errors. Every system trades off missed detections against false alarms. That trade-off is a business decision, not a technical one.

That is why good systems keep a human in the loop where errors are expensive: the model points, the person decides. That is how well-designed document capture works and how systems that support radiologists work.

What about privacy and the AI Act?

An image in which a person can be identified is personal data under GDPR, and biometric data used for identification is a special category (Article 9 GDPR) whose processing is prohibited except under specific exceptions (GDPR).

The EU AI Act goes further. Since February 2, 2025 it has banned, among other things, emotion recognition in the workplace and in education (except for medical or safety reasons) and building facial recognition databases through untargeted scraping of images from the internet or CCTV footage (AI Act, Article 5). Biometric systems listed in Annex III are high-risk; after the amendment by Regulation (EU) 2026/1744, their obligations apply from December 2, 2027 (Regulation (EU) 2026/1744).

Most industrial applications (defects, pallets, document capture) do not involve people and do not trigger these rules. When a camera covers employees or customers, start with a legal review. Local processing also helps, because images never leave your company. We show what AI on your own infrastructure looks like on our AI infrastructure page.

How do you start a computer vision project?

With one well-defined problem and images from the real environment. An order that works in practice:

  1. Define the decision. What the system must decide and what an error costs in each direction.
  2. Collect a sample from real conditions. Different shifts, cameras and lighting, and above all the hard cases.
  3. Write down labeling criteria. One shared definition of "good" and "bad".
  4. Measure the baseline. How much time and which errors today's manual process produces.
  5. Decide where the model runs. Cloud, on-premise server or a device next to the camera, depending on response time and data requirements.
  6. Keep a human where errors are expensive, and collect their corrections as new training data.

Computer vision is no longer an experiment for the largest companies. It is a tool for very specific work: looking at the things people look at by hand today, faster and to a consistent standard.

Sources