Deep MIB - 2D Instance Segmentation (SOLOv2)¶
Training and application of the SOLOv2 network for 2D instance segmentation in Microscopy Image Browser.
Overview¶
Unlike the semantic workflows, which assign every pixel to a class (all objects of the same type share one label), instance segmentation detects each object individually: every object receives its own distinct label, so touching or overlapping objects of the same type are separated. Deep MIB implements 2D instance segmentation using the MATLAB SOLOv2 network.
The typical procedure follows the same three phases as the other workflows — with one important difference: instance segmentation requires a preprocessing step before training (see Preprocessing).
- Preprocess the ground-truth models into the SOLOv2 annotation format (Directories and Preprocessing tab).
- Train the SOLOv2 network (Train tab).
- Predict new datasets and export the detected objects as MIB models (Predict tab).
Support package required
solov2 requires the Computer Vision Toolbox Model for SOLO V2 Instance Segmentation
support package. To install this support package, use the Add-On Explorer.
Compatible source datasets¶
- Images: 2D images with 1 (grayscale) or 3 (RGB) colour channels. Grayscale images are automatically converted to RGB, because the SOLOv2 backbone expects 3-channel input.
- Ground-truth labels: a MIB
.modelfile per image (or a single.modelfile for all images when is checked), in which each individual object is painted with its own unique index (1, 2, 3, …) and the background is0. Touching objects must use different indices to be separated.uint8models allow up to 255 objects per image;uint16models allow many more. - Classes: the current implementation is single-class — every detected object belongs to
one class named
object. The field is not used for this workflow.
Preparing instance labels
Paint each object as a separate material index in the MIB model (for example, using the
watershed / object-separation tools), so that regionprops can extract one bounding box and
one mask per object during preprocessing.
Network panel¶
In the Network panel select:
- : 2D Instance
- : SOLOv2
- : Resnet18 (lighter, faster) or
Resnet50 (heavier, potentially more accurate). These map to the pretrained
light-resnet18-cocoandresnet50-cocobackbones.
The is given as
height width depth colors, for example 800 800 1 3. The width and height should be
multiples of 32, the depth is always 1, and the colour value is 3. This size defines the
patch cropped during training and the tile used during prediction — images are processed
at native resolution in patches/tiles of this size, not downscaled as a whole.
Note
Use a square input patch size if you intend to enable the 90° rotation augmentations (see the Train tab); non-square patches are not compatible with those augmentations.
Preprocessing (required for training)¶
Instance segmentation does not train directly on MIB .model files: preprocessing first
converts each model into a lightweight MAT-file that stores the 2-D instance label map.
Training then crops native-resolution patches from these maps on-the-fly.
To preprocess for training:
- Arrange the training data under the training directory (named
1_Trainingin the directory schemes):Images/— the training imagesLabels/— the corresponding.modelfiles (one per image, or a single model file)
- Set to Training.
- Press . Deep MIB converts the models into
LabelsInstances/*.matand then offers to split the files into training and validation sets. Choosing Split labels creates theTrainImages,TrainLabels,ValidationImagesandValidationLabelssubfolders used during training.
Preprocessing writes one .mat per input image (so the file count matches the number of
input models), splitting them into the TrainLabels/ValidationLabels folders.
Preprocessed label file format (LabelsInstances/*.mat)
Each MAT-file is lightweight and contains only two variables:
| Variable | Type | Description |
|---|---|---|
imageFilename |
char |
name of the source image (e.g. 00001.png), used to locate the image in TrainImages / ValidationImages |
instanceLabelMap |
H×W uint16 |
2-D label map — each object a unique index, background 0 |
Storing a compact 2-D label map (instead of a full H×W×N binary mask stack) keeps memory
usage low for large / whole-slide images. Native-resolution training patches are cropped from
this map on-the-fly, and the per-patch object masks and bounding boxes are rebuilt only for the
small cropped patch.
Prediction needs no preprocessing
Only training requires preprocessing. For prediction, SOLOv2 reads the raw images
directly, so no label preprocessing is performed — place the images to segment in the
prediction directory (2_Prediction/Images) and run the Predict tab.
Training¶
Training is configured and started from the Train tab.
- Native-resolution patches: the network trains on patches cropped on-the-fly from the
label maps at native resolution (the network input size, e.g.
800×800), not on downscaled whole images. This is essential for large / whole-slide microscopy images, where resizing the whole image to the network input would destroy resolution and hide small objects. Images smaller than the patch are symmetrically padded. - Patch sampling: each image contributes patches per epoch. ~90 % are object-seeded (a random object is picked and the patch placed with a random offset so the object appears anywhere in the patch, never forced to the centre) and ~10 % are uniform-random (may be pure background).
- Augmentation (): supported for instance segmentation. Geometric augmentations (reflections, 90°/arbitrary rotation, scale, shear) are applied to the image and every object mask together, and the bounding boxes are recomputed from the transformed masks so that boxes, masks and labels stay in sync. Intensity/colour augmentations (noise, blur, brightness/contrast/hue/saturation jitter) are applied to the image only. Augmentation settings are shared with the 2D semantic workflow.
- Validation: a validation set (created during the split step) is used when available; if the installed MATLAB version does not support validation for SOLOv2, training automatically retries without validation.
- Progress: the SOLOv2 trainer reports the training loss, shown in the training progress
window. There is no training accuracy metric for this workflow (the Training accuracy gauge
reads
N/A), but an optional validation accuracy can be enabled — see Validation accuracy (mAP) below.
Validation accuracy (mAP)¶
By default the trainer computes only the validation loss. To also track a validation accuracy, enable on the Options tab. When checked, Deep MIB attaches the mean Average Precision (mAP) metric to the SOLOv2 training: at each validation interval the trained detector is run over the whole validation set and the resulting mAP is displayed on the Validation accuracy gauge of the custom training progress window.
- Requires a validation set (Directories and Preprocessing → Fraction of images for validation > 0);
with no validation images the metric cannot be computed and the gauge stays
N/A. - It adds a noticeable cost — full inference plus IoU matching over the validation set on every validation pass — so leave it unchecked to speed up training when only the validation loss is needed.
Training parameters for the 2D Instance workflow
| Parameter | Where | Notes / recommendation |
|---|---|---|
| Input patch size | Train tab | height width 1 3; width/height multiples of 32; use a square size to allow 90° rotation augmentations. Patches are cropped at this size from the images at native resolution. |
| Encoder | Network panel | Resnet18 (lighter, faster) or Resnet50 (heavier, potentially more accurate). |
| Patches per image | Train tab | number of native-resolution patches sampled from each image per epoch; increase for large images that hold many objects. |
| Mini-batch size | Train tab | patches processed per iteration; bounded by GPU memory. Must be ≤ Patches per image × number of training images. |
| Augmentation | Train tab | enable to apply the shared 2D augmentations (geometric to image + masks, intensity to image only). |
| Number of epochs | Train settings | more epochs = more augmented exposure (each epoch re-crops and re-augments every image). |
| Solver / Initial learn rate / schedule | Train settings | e.g. sgdm, initial learn rate 5e-4, piecewise schedule (as in the MathWorks SOLOv2 example); adam/rmsprop also supported. |
| Fraction of images for validation | Directories and Preprocessing tab | fraction moved to the validation set during the split; 0 disables validation. |
| Random generator seed | Directories and Preprocessing / Train tab | fix for reproducible splits and training; 0 = new random seed each run. |
Only training uses patch cropping; Patches per image has no effect on prediction (which tiles whole images — see below).
The trained network is saved to the *.mibDeep file specified by
, together with its configuration
(*.mibCfg).
Prediction¶
Prediction is started from the Predict tab:
- Set to the trained
*.mibDeepfile. - Place the images to segment in the prediction directory (
2_Prediction/Images). - Press .
Large images are tiled at native resolution using the blocked-image overlap strategy
(controlled by and
in the Predict tab): each tile is segmented with
segmentObjects with a surrounding border of context, and objects are stitched across tiles. The
stitching rule is selected with the
dropdown of the Instance segmentation subpanel in the Predict tab:
- Centroid in core — each object is emitted by the tile that owns its centroid, so there are no duplicates and no seam-splitting. Requires the tile overlap to be at least as large as the biggest object.
- IoU merge — all detections of every tile are kept, and detections from neighbouring tiles are merged into one instance when their masks agree inside the shared overlap band (intersection-over-union test). Objects larger than the overlap are detected piecewise and merged, so only the band width matters — use this mode when objects may exceed the overlap.
The detection confidence threshold, the merge IoU/IoA thresholds and the minimal object area are configured with the subpanel's Settings button — see the Predict tab for details.
The result is written as a MIB .model file under
3_Results/PredictionImages/ResultsModels, in which every object instance is a unique index
(background 0), matching the input labelling convention. Two objects that are separated in the
image never share an index: in both overlap modes each index is finally split into its separate
objects, and components below
are discarded.
Both 2D images and z-stacks may be used as prediction images, exactly as in the 2D Semantic workflow:
- a 2D image is predicted directly and saved as a 2D model;
- a z-stack (a multi-page TIF, an HDF5 volume, ...) is predicted slice-by-slice and saved as a single 3D model with the same number of slices.
In both cases the instance indices run 1..N within each slice and are not consistent
between slices - an object continuing through several slices carries a different index on each of
them. Linking them into true 3D objects is done afterwards with
Merge 2D to 3D, which ignores the input indices.
Overlap size vs stitching mode
With Centroid in core, the tile overlap must be at least as large as the biggest object — otherwise an object that never fits fully inside a single tile's field of view will be truncated. Increase for large objects, or switch to IoU merge, which lifts this restriction (the overlap band only needs to be wide enough — a few tens of pixels — for neighbouring tiles to produce consistent masks in it). If IoU merge is selected while is unchecked, a default 5% overlap is applied automatically.
Merging 2D predictions into a 3D model¶
When the prediction images are serial sections of a volume, the predicted instance models can be merged into 3D instance models with the button of the Instance segmentation subpanel in the Predict tab. Objects overlapping between neighbouring slices are linked into 3D instances with one consistent index through the whole stack.
The layout of the results in 3_Results/PredictionImages/ResultsModels is detected automatically
from the depth of the first *.model file:
- 2D models (the prediction images were separate 2D files) - all files are the Z-slices of one stack and are taken in alphabetical order of their filenames, so make sure the prediction images are named in their correct Z-order. One merged 3D model is written.
- 3D models (the prediction images were z-stacks) - every file already is a complete stack, so each is stitched independently and one merged 3D model is written per input file.
Do not mix the two layouts
The results folder must contain either 2D models only or 3D models only. A mixed folder is ambiguous (are the 2D files slices of their own stack, or strays?) and is rejected with an error. Empty the folder before re-running a prediction on a different kind of input.
The stitching settings dialog is the same one used by
Ribbon → Model → Stitch 2D instances to 3D,
where every option is described and illustrated. One option differs here: Z anisotropy ratio is
a number to type in rather than a checkbox, because raw prediction images carry no pixel size for it
to be taken from (1 = isotropic, off).
The dialog reopens on the values used last for as long as MIB is running, and the two entry points share them, so a threshold tried in one is offered in the other. Each accepted run also prints one line to the MATLAB console listing exactly what was used - handy when trialling thresholds over several runs. The values are per session and are not written to preferences.
After the settings, the destination is requested. When a single merged model is produced, a
file dialog asks for the directory, filename and file format. When several stacks are stitched
(one per 3D input model), a folder is requested instead, followed by a format dropdown; each output
is then named after its input model with a _stitched3D suffix.
The merged model can be written as a single 3D file (e.g. Matlab format (*.model), TIF 3D stack, Amira Mesh, HDF5, MRC, NRRD) or as a sequence of 2D files (e.g. Matlab format 2D sequence (*.model), TIF/PNG 2D sequence) — for TIF the policy is asked during saving of a single model, while multiple models are always written with the format's default policy to avoid one dialog per file.
Merging takes a while. It can be stopped at any point with ; when several stacks are being merged, the models finished before the stop are kept and the rest are not written.