Image Segmentation Explained: Semantic, Instance and Panoptic Segmentation

What is image segmentation

Image segmentation identifies regions within an image, often represented as pixel masks. Depending on the task, a mask can show a category such as grass, separate individual dogs, or mark a selected region without naming its class.

A mask is an image-aligned map: its positions correspond to locations in the image. A binary mask marks inside versus outside a region. Other outputs attach class labels or object identities. Overlay colors help us see those assignments; they are display choices.

Segmentation is a broad family of tasks. This guide focuses on common AI-based approaches, using one scene throughout. The important question is what each output represents, rather than which color is used to draw it.

This branch of computer vision is useful when shape matters. To cut a dog out of a photograph, knowing that the image contains dogs is insufficient. We need the region belonging to the chosen dog.

Segmentation compared with classification and detection

Imagine two dogs standing on grass beneath a sky. Image classification describes the image with labels. Object detection locates objects with boxes. Segmentation describes pixel regions.

The same two-dog image is described with an image label, bounding boxes and object masks.
Labels boxes and masks

Labels describe an image, boxes locate objects and masks describe regions.

Dogs is one image-level label, not a label placed on a box. Classification can also be multi-label. The detection panel contains two boxes; the mask panel shows two object regions. Boxes often include background, and predicted masks can still be wrong.

View the diagram at full size

A box tells an editor where to look; a mask gives a region to select. These outputs can coexist: Mask R-CNN predicts object boxes and instance masks. Segmentation does not always require a box first.

If the goal is to change only one dog’s appearance, the selection should follow that dog’s outline. A rectangle may include grass between its legs. A predicted mask can reduce that unwanted selection, but still needs checking.

Semantic instance and panoptic segmentation

Semantic segmentation

Semantic segmentation assigns class labels to pixels under a defined label scheme. In our scene, both dogs share the Dog class, even though their regions are disconnected. Grass and sky have their own labels. The result does not identify two separate dogs.

Instance segmentation

An instance is an individual object. Instance segmentation separates Dog 1 from Dog 2, typically with a mask for each. Our object-only example leaves grass and sky unsegmented. That neutral area should not be read as a predicted background class.

Panoptic segmentation

Panoptic segmentation combines class labels with separate identities for countable objects in one non-overlapping map. The Panoptic Segmentation paper distinguishes countable “things,” such as our dogs, from “stuff,” such as grass and sky. Dataset definitions determine the division; ambiguous or excluded pixels may be void or ignored.

Semantic masks share a dog class, instance masks separate two dogs, and panoptic masks also label grass and sky.
Three segmentation tasks in one scene

Class and object identity answer different questions.

Semantic: both dogs share the Dog class and have no separate IDs. Instance: Dog 1 and Dog 2 are separate; the neutral grass and sky are Not segmented, not a predicted background class. Panoptic: both dog IDs plus grass and sky appear in one non-overlapping map. Dogs are things; grass and sky are stuff in this example. Dataset rules may allow void or ignored pixels.

View the diagram at full size

TaskClass regions shownSeparate dog identities
SemanticDog, grass and skyNo
InstanceDog objects onlyYes
PanopticDog, grass and skyYes

How a model produces a mask

Learn from annotated examples

Supervised training commonly pairs images with reference masks, as the TensorFlow tutorial demonstrates. Training adjusts model parameters to reduce prediction errors. Transfer learning can reuse pretrained representations.

Predict and interpret the output

A trained model extracts features and predicts masks. Convolutional neural networks are one approach; others use transformers. The original Segment Anything Model, or SAM, can use clicks or boxes to guide a mask. Guidance can identify a desired region without assigning a semantic class.

Using the trained model is called inference. Ordinary inference applies learned parameters; it does not automatically retrain the model on the new image. The output must be interpreted according to the task and label scheme.

An image is processed into features and a predicted mask that is then reviewed.
From image to predicted mask

A trained model predicts a mask; some systems also accept guidance such as a click.

Training on examples happens before this prediction workflow. The diagram does not show incoming images updating model weights. Prompt optional means that some systems accept a click or other guidance; a prompt may be ambiguous. Features are an abstract schematic, not measured activations.

View the diagram at full size

When segmentation is useful

Choose the output around the job. An editor may need only a selected subject. A road-scene analysis may need class regions such as road and sidewalk. For segmentation-based counting, separate instance masks make individual objects countable. Biomedical research such as V-Net also studies anatomical outlines, with application-specific expert review.

Ask what must be covered, which objects must remain distinct and which mistakes matter. A mask alone establishes neither a diagnosis nor a safe driving decision. Counting its pixels also does not establish physical area without suitable scale and geometry. The useful output depends on the application and its validation.

How to measure segmentation quality

A reference annotation is the labeled mask used for comparison, often called “ground truth.” It follows a labeling policy and is not infallible.

Foreground means the selected region. Foreground intersection over union, or IoU, divides the shared foreground area by the area marked foreground in either mask. Our separate grid makes the calculation visible.

Two eight-cell foreground masks share six cells and cover ten cells in total, giving an IoU of 60 percent.
Count the overlap

Foreground IoU counts shared cells divided by the union, here 6 out of 10.

This is a separate 4 × 5 toy grid, not a rasterized dog scene. Reference and Prediction each have 8 foreground cells. Solid blue: 6 shared. Diagonal hatching: 2 reference-only, missed cells. Dots: 2 prediction-only, extra cells. White: 10 neither. Union = 8 + 8 − 6 = 10. Foreground IoU = 6 ÷ 10 = 0.60 = 60%; this is not pixel accuracy or an overall model score.

View the diagram at full size

In another example, 95 of 100 pixels are background. Predicting background everywhere gives 95% pixel accuracy but zero foreground IoU: all five foreground pixels were missed.

Choose metrics for the task. Semantic evaluation commonly uses class-wise IoU and mean IoU, which averages class results. V-Net illustrates another overlap measure, Dice. Instance mask average precision accounts for object matching and confidence rankings; panoptic quality combines matched-region overlap with recognition errors. The Cityscapes benchmark rules show why dataset, averaging, matching and ignored/background rules must accompany scores.

For these masks, Dice doubles the shared area and divides by both foreground counts combined: 12 divided by 16, or 75%. Metric percentages are not interchangeable.

Why segmentation masks can be wrong

Boundaries and hidden regions

Thin tails, low contrast and touching objects can make boundaries difficult. Ordinary region overlap can underrepresent edge errors, especially for large objects, as Boundary IoU explains. Inspect edges alongside scores.

A predicted mask misses a thin tail, while another incorrectly merges two dog instances.
Boundary and instance errors

Boundary quality and object separation are different checks.

Deliberately altered teaching examples. The circles compare the same tail location: Prediction misses the tail and has a jagged back edge; Reference retains the tail. The touching-dog prediction merges two instances into one mask, while Reference keeps IDs 1 and 2. This is an instance-separation error, not an error merely because semantic dog pixels share a class. Reference annotations follow a labeling policy; fur and ambiguous boundaries can admit different conventions. Visible-region masks do not reconstruct hidden parts.

View the diagram at full size

Occlusion means part of an object is hidden. Standard visible-region masks cover what is visible; reconstructing an entire hidden shape belongs to a separate task called amodal segmentation, distinguished in the panoptic paper.

Labels and unfamiliar images

The Cityscapes labeling policy shows why annotation conventions matter. People may disagree about fur, shadows or unclear boundaries. Those choices affect both training targets and evaluation.

Different lighting, cameras or environments can also change performance, a problem called domain shift. Check wrong classes, missed objects, merged instances, split instances and inaccurate edges separately. Shared class labels alone are expected in semantic segmentation. A colorful mask or confident score does not guarantee correctness.

Evaluate held-out images that reflect intended use. Inspect small or rare regions separately, since a useful overall result can still hide repeated failures on the details a particular application needs.

Frequently asked questions

What is a segmentation mask?

A segmentation mask maps region membership or labels to image locations. A binary mask marks inside versus outside. Display colors make regions visible; the mask does not necessarily contain a class name.

How is segmentation different from object detection?

Object detection usually locates an object with a rectangular box. Segmentation describes its pixel region and shape. Some systems produce both outputs, and neither guarantees an exact object boundary.

What is the difference between semantic and instance segmentation?

Semantic segmentation groups pixels by class. Instance segmentation separates individual objects, so two dogs can have separate masks while sharing the same dog category. Semantic output alone supplies no separate dog identities.

What does panoptic segmentation add?

Panoptic segmentation combines scene-level class labels with separate identities for countable objects. Its final segments do not overlap. Dataset rules may still reserve void or ignored pixels for ambiguous or excluded regions.

Does every segmentation model label every pixel?

No. Object-only and prompted tasks can return selected masks while leaving other regions unlabeled. Dense semantic tasks aim for per-pixel classes within their defined label scheme, subject to the dataset’s rules.

Can segmentation work from a click or a box?

Yes. Promptable systems such as the original SAM can use clicks or boxes as guidance. A prompt may refer to several plausible regions, so candidate masks still need inspection and sometimes further guidance.

What is IoU in image segmentation?

Intersection over union is the shared foreground area divided by the area covered by either mask. Specify the class or object being compared, then explain how results across examples or classes are aggregated.

Why do masks miss edges or merge objects?

Fine structures, low contrast, occlusion and ambiguous boundaries make prediction difficult. Training examples and annotation rules influence what gets included or separated. Merging two objects is an error when separate instances are required.

Does a high segmentation score mean the model is ready to use?

No. Check relevant classes, small objects, boundaries and unfamiliar conditions. Consider error consequences and application requirements. A single overlap score cannot establish readiness, and a good average can hide important failures.

Key takeaway and next step

A mask describes a region. Semantic segmentation adds class labels; instance segmentation separates objects; panoptic segmentation brings both into a scene map. Start by asking which output your task needs, then inspect the mistakes that matter. Read Model Evaluation Metrics Explained for broader evaluation principles, keeping segmentation’s task-specific matching and boundary checks in view.

Scroll to Top