Reinnder LabsNow · 2026 to 2027

How robots see: computer vision and depth sensing, explained

Robot vision is the combination of cameras or depth sensors with software that finds objects, works out where they are in three dimensions and lets a robot reach, pick or avoid them.

Now Reinnder Robotics Updated 4 October 2026 · 4 min read

Stage
Now · 2026 to 2027
Area of work
Reinnder Robotics
Used in
Bin picking · Quality inspection
Key terms
4 defined · glossary

Figure · Illustration

From a camera image to a grasp

  1. Cameras or depth sensors record the scene, with the lighting kept as steady as possible.

  2. Software corrects lens distortion and removes noise and glare from the image.

  3. A trained model finds each part and, where needed, its outline.

  4. The part’s position and angle are worked out and converted into the robot’s own coordinates.

  5. The robot chooses a grasp that avoids collisions, moves, and then checks that the grasp worked.

Illustration of a typical vision-guided pick. Real systems differ in their sensors and software.

From pixels to a position#

A camera produces a grid of pixels, not objects. Software has to find the object in the picture, usually with a neural network trained on many labelled examples, then work out where it is in the room. For a robot arm that means converting a position in the image into a position in the robot’s own coordinates, using a calibration done once, carefully, when the camera is mounted. Without that step the robot sees the part and misses it.

Seeing depth#

A single ordinary camera does not measure distance. Robots get depth in several ways: two cameras side by side, which compare their two views as eyes do; a sensor that projects a pattern or a pulse of infrared light and measures what returns; or lidar, which times laser pulses. Each has strengths: stereo works in daylight, projected patterns are accurate at short range indoors, and lidar covers long distances.

Why it is harder than it looks#

A factory is not a photograph studio. Shiny metal makes reflections, a dark part can vanish, a pile of parts hides most of what is under it, and the lighting changes through the day. Systems are therefore trained on examples that include the awkward cases, tested on real parts in the real light, and often paired with a second sensor, such as force sensing in the gripper, to confirm that the grasp worked.

Where it is used#

  • Bin picking Finding a part in a pile and choosing a way to grasp it.
  • Quality inspection Spotting a scratch, a missing part or a wrong label.
  • Warehouses Recognising boxes and parcels of different sizes.
  • Farms Telling a ripe fruit from an unripe one before picking.

What has made robot vision practical?#

Three changes. Learned detection models now find objects in cluttered images far more reliably than hand-written rules did. Depth cameras have become small and affordable. And graphics processors and efficient chips can run these models fast enough at the machine. Together they moved vision from a specialist add-on, needing an expert for every new part, to something a small team can set up.

The remaining difficulty is less the algorithm than the conditions: light, reflections, dirt and the variety of real parts.

Robot vision compared with fixed positioning#

A traditional robot repeats the same motion to the same point and relies on fixtures to hold parts exactly in place. Vision lets a robot adapt, at the price of extra cost, lighting control and tuning.

If parts always arrive in the same place, a fixture is simpler and more reliable. Vision pays when parts vary, arrive loose or change often.

Three ways a robot can get depth, compared
Compared onSingle cameraStereo pairDepth sensor
How it gets depthEstimates it by learningCompares two viewsMeasures it directly with light
CostLowestLow to moderateModerate
StrengthCheap, with rich detailWorks outdoors in daylightReliable at short range indoors
WeaknessScale is partly a guessStruggles on blank surfacesCan be washed out by sunlight

What should you test before buying a system?#

Insist on a trial with your own parts, not a demonstration with the supplier’s. Include the worst cases: the shiniest part, the darkest, a full bin, a part lying on its side. Run it in the lighting of your site, at different times of day, and for hours, not minutes. Ask how long it takes to teach a new part, what happens when it is unsure, and whether it reports how many picks failed.

A good system fails safely: it stops or asks for help rather than guessing.

Key terms#

Camera calibration
Measuring where a camera sits relative to a robot, so image positions can be turned into robot positions.
Point cloud
A set of points in three dimensions, each marking where a sensor found a surface.
Object detection
Software that finds and labels objects in an image.
Pose estimation
Working out an object’s position and orientation, not just that it is there.

Common questions#

Does a robot need a 3D camera to pick things?

Not always. Flat parts on a flat surface can be handled with a single camera. 3D sensing helps when parts are piled, tilted or vary in height.

What is a vision-guided robot?

A robot whose movements are corrected by what a camera sees, so it can handle parts that are not in exactly the same place each time.

Does AI remove the need for camera calibration?

No. A robot still has to know where its camera sits relative to its arm. Learned methods can ease the process but not skip the question.

Sources and further reading#

Independent pages we checked while writing this guide. They are not Reinnder products, and Reinnder is not affiliated with them.

Reinnder’s angle

How Reinnder looks at how robots see

Reinnder Robotics is planned to include vision that lets a work cell cope with parts that are not always in the same place.

Reinnder Robotics

This guide explains the technology in general terms. It is not advice, and it does not describe a Reinnder product on sale.