Log in Sign up
Back to Discover
💻

Computer vision

technology Maturity 7-9

Computers can see things too.

Intersection over Union - object detection bounding boxes.jpg
Intersection over Union - object detection bounding boxes.jpg
They look at pictures. They can tell what is in a photo. This helps robots move around. It helps machines work. It is very cool! Can you see things with your eyes?

38 words

Computers can learn to see.

Intersection over Union - object detection bounding boxes.jpg
Intersection over Union - object detection bounding boxes.jpg
They look at pictures or videos. They can find things in a photo. This helps them know what is there.

They can even see in 3D. This helps them understand shapes.

Synthesizing 3D Shapes via Modeling Multi-View Depth Maps and Silhouettes With Deep Generative Networks.png
Synthesizing 3D Shapes via Modeling Multi-View Depth Maps and Silhouettes With Deep Generative Networks.png
A computer can look at a flat picture. Then it can build a model of the world.

This helps robots move through a room. It helps machines work in a factory. It can even help with medical scans.

LiDAR Scanner and Back Camera of iPad Pro 2020 - 3.jpg
LiDAR Scanner and Back Camera of iPad Pro 2020 - 3.jpg
Computers are learning to see like we do!

98 words

Computers can learn to see. This is called computer vision.

Intersection over Union - object detection bounding boxes.jpg
Intersection over Union - object detection bounding boxes.jpg
It is a way for machines to understand pictures and videos. The goal is to make computers see like humans do. This helps them make smart choices based on what they see.

Computers use many kinds of data. They can look at video clips. They can use 3D scanners. Some tools use LiDaR, which is a way to map things.

LiDAR Scanner and Back Camera of iPad Pro 2020 - 3.jpg
LiDAR Scanner and Back Camera of iPad Pro 2020 - 3.jpg
This helps them see shapes and depth. They can even find objects in a photo. They can track how things move.

Scientists study how to do this. They use math and physics. They also look at how brains work. This is called neurobiology.

Synthesizing 3D Shapes via Modeling Multi-View Depth Maps and Silhouettes With Deep Generative Networks.png
Synthesizing 3D Shapes via Modeling Multi-View Depth Maps and Silhouettes With Deep Generative Networks.png
They build models that act like eyes and brains. New ways of learning, called deep learning, make computers very good at this. They can now recognize faces and objects with great skill. This helps robots move and helps doctors with medical scans.

171 words

Computer vision is a special field of science. It helps computers understand digital images and videos.

Intersection over Union - object detection bounding boxes.jpg
Intersection over Union - object detection bounding boxes.jpg
Most computers just see a grid of colors. Computer vision lets them see the world instead. It turns pictures into descriptions that make sense. This allows a machine to make smart decisions. It can tell what an object is or where it is. This is very important for many modern technologies.
DARPA Visual Media Reasoning Concept Video.ogv
DARPA Visual Media Reasoning Concept Video.ogv

How does this work step by step? First, the computer must acquire the image data. It might use a regular camera or a 3D scanner. Some systems use LiDaR sensors to see depth.

LiDAR Scanner and Back Camera of iPad Pro 2020 - 3.jpg
LiDAR Scanner and Back Camera of iPad Pro 2020 - 3.jpg
Next, the computer processes the visual information. It uses math and physics to look at shapes. It might look for edges or lines in a picture. Then, it analyzes the data to understand the scene. Finally, it extracts useful information to act on. This might mean tracking a moving object in a video.

People have been studying this since the late 1960s. It started at universities that studied artificial intelligence. In 1966, researchers thought it was a simple task. They believed a summer project could solve it. They wanted to attach a camera to a computer. The goal was to make it describe what it saw. This was meant to help robots act intelligently. In the 1970s, scientists built the first real foundations. They learned how to find edges and motion.

Synthesizing 3D Shapes via Modeling Multi-View Depth Maps and Silhouettes With Deep Generative Networks.png
Synthesizing 3D Shapes via Modeling Multi-View Depth Maps and Silhouettes With Deep Generative Networks.png

Many different tools and numbers help this field grow. In the 1990s, scientists used math to recognize faces. This used a method called Eigenface. Later, researchers used graph cuts for image segmentation. This helps a computer separate an object from its background. Today, we use deep learning to make systems better. Deep learning uses complex math to learn from data. It is much more accurate than older methods. This helps with tasks like classification or motion estimation.

SRT Shape Recognition Technology.png
SRT Shape Recognition Technology.png

Computer vision links to many things you might know. It is closely tied to neurobiology, which is the study of brains. Scientists look at how eyes and neurons work. They use these ideas to build artificial systems. It also uses physics to understand light and sensors. Even robots use it to find their way. A robot needs vision to plan a path safely. You might even see it used in fashion or medicine. It helps doctors look at medical scans more closely. It makes our world much more automated and smart.

427 words

Computer vision is an interdisciplinary field focused on how computers gain high-level understanding from digital images or videos. At its core, it involves the automatic extraction, analysis, and understanding of useful information from visual data. This process transforms raw visual images into descriptions of the world that make sense to thought processes. These descriptions can then elicit appropriate actions from a machine. From an engineering perspective, the goal is to automate tasks that the human visual system performs naturally.

Intersection over Union - object detection bounding boxes.jpg
Intersection over Union - object detection bounding boxes.jpg

The mechanism of computer vision begins with the acquisition of data. This data can take many forms, such as video sequences, views from multiple cameras, or multi-dimensional data from a medical scanner. Some systems use LiDaR sensors to collect 3D point clouds.

LiDAR Scanner and Back Camera of iPad Pro 2020 - 3.jpg
LiDAR Scanner and Back Camera of iPad Pro 2020 - 3.jpg
Once the data is acquired, the system must process and analyze it. This involves disentangling symbolic information from image data. Scientists use models constructed with the aid of geometry, physics, statistics, and learning theory to achieve this. The ultimate result is the production of numerical or symbolic information, which allows the computer to make decisions.

There are many specialized subdisciplines within computer vision that handle different tasks. Scene reconstruction and 3D scene modeling focus on rebuilding the structure of a space. Object detection and object recognition allow a system to identify specific items. Event detection and activity recognition focus on understanding what is happening over time. Other tasks include video tracking, 3D pose estimation, and motion estimation. There are also technical processes like image restoration, indexing, and visual servoing. Each of these subfields applies different mathematical and algorithmic approaches to solve specific visual problems.

The history of computer vision began in the late 1960s at universities pioneering artificial intelligence. Researchers wanted to mimic the human visual system to help robots behave intelligently. In 1966, it was famously believed that a simple undergraduate summer project could solve the problem. The plan was to attach a camera to a computer and have it "describe what it saw." However, the field proved much more complex than that initial goal. During the 1970s, researchers built foundations by studying edge extraction, line labeling, and motion estimation.

Synthesizing 3D Shapes via Modeling Multi-View Depth Maps and Silhouettes With Deep Generative Networks.png
Synthesizing 3D Shapes via Modeling Multi-View Depth Maps and Silhouettes With Deep Generative Networks.png

As the field progressed, it became more mathematically rigorous. The 1980s saw the introduction of concepts like scale-space and the inference of shape from shading or texture. In the 1990s, research into projective 3-D reconstructions improved camera calibration. This era also saw the first use of statistical learning techniques to recognize faces, a method known as Eigenface. Toward the end of the 1990s, computer vision began to interact more closely with computer graphics. This led to new techniques like image morphing, panoramic image stitching, and view interpolation.

DARPA Visual Media Reasoning Concept Video.ogv
DARPA Visual Media Reasoning Concept Video.ogv

Today, the field is experiencing a resurgence due to the advancement of deep learning. Deep learning is a type of machine learning that uses complex optimization frameworks. These algorithms have surpassed prior methods in accuracy for tasks like classification, segmentation, and optical flow. This technological leap allows for much more sophisticated artificial vision systems. The field also maintains a distinction from machine vision, which is a systems engineering discipline used mostly in factory automation. While the terms have converged recently, they represent different focuses within the broader industry.

Computer vision is deeply connected to several other scientific fields. Neurobiology has a massive influence, as researchers study how eyes, neurons, and brain structures process visual stimuli. An early example of this influence is the Neocognitron, a neural network developed by Kunihiko Fukushima in the 1970s. The field also relies on solid-state physics to design image sensors that detect electromagnetic radiation. Furthermore, computer vision is essential for robotic navigation. A robot requires a detailed understanding of its environment to perform autonomous path planning. This makes computer vision a vital component of modern robotics, medicine, and even fashion eCommerce.

649 words
🖼️ Images & Media (7)
File:Intersection_over_Union_-_object_detection_bounding_boxes.jpg
Intersection_over_Union_-_object_detection...
File:Synthesizing 3D Shapes via Modeling Multi-View Depth Maps and Silhouettes With Deep Generative Networks.png
Synthesizing 3D Shapes via Modeling...
DARPA Visual Media Reasoning Concept Video.ogv
File:Mars Science Laboratory, 2011-Present.jpg
Mars Science Laboratory, 2011-Present.jpg
Finger sensor.webp
File:SRT Shape Recognition Technology.png
SRT Shape Recognition Technology.png
File:LiDAR_Scanner_and_Back_Camera_of_iPad_Pro_2020_-_3.jpg
LiDAR_Scanner_and_Back_Camera_of_iPad_Pro_...
Up Next
💻
Deep learning
Technology
More to explore

What is Nepedia?

A free, ad-free encyclopedia for children. Every article is written at five reading levels, so the same page works for a five-year-old and a fifteen-year-old — use the level switcher above to see this one change. No account needed to read.