More in Computer Vision — page 5
Design a tracking-by-detection tracker
WHAT IT TESTS: building tracking from detection plus data association. OUTLINE: detect per frame, then associate boxes across frames by IoU or appearance using Hungarian matching, maintaining track ids.
Sparse vs dense optical flow and Lucas-Kanade
WHAT IT TESTS: understanding motion estimation granularity. OUTLINE: sparse flow tracks selected feature points, dense flow computes a vector per pixel; Lucas-Kanade solves brightness constancy in a local window assuming constant motion.
Adapting ViT for dense semantic segmentation
WHAT IT TESTS: turning a classification ViT into a dense predictor. OUTLINE: reassemble patch tokens into a 2D feature map, add a decoder, and handle low resolution plus quadratic attention cost.
Explain panoptic segmentation and Panoptic Quality
WHAT IT TESTS: unifying semantic and instance segmentation plus its metric. OUTLINE: panoptic assigns every pixel a class and instance id over things and stuff; PQ factors into SQ, average IoU of matches, times RQ, an F1 over matched segments.
How to improve coarse segmentation boundaries?
WHAT IT TESTS: practical debugging of low-resolution mask edges. OUTLINE: skip connections and higher-resolution features, boundary-aware losses, and point-based or CRF refinement.
How does Mask R-CNN do instance segmentation?
WHAT IT TESTS: understanding of two-stage detectors and per-instance masks. OUTLINE: Faster R-CNN backbone plus RPN, then RoIAlign and a parallel mask head predicting per-class binary masks. RED FLAG: claiming masks are shared or that RoIPool is used.
U-Net architecture and its skip connections
WHAT IT TESTS: encoder-decoder design for segmentation. OUTLINE: U-Net has a contracting encoder, an expanding decoder, and skip connections that concatenate matching-resolution encoder features into the decoder to recover spatial detail lost in downsampling.
Semantic versus instance segmentation
WHAT IT TESTS: distinguishing two pixel-labeling tasks. OUTLINE: semantic segmentation labels each pixel by class but merges objects of the same class; instance segmentation also separates individual objects.
Detector head losses: regression versus classification
WHAT IT TESTS: multi-task loss design in detection heads. OUTLINE: the head splits into a classification branch using cross-entropy over classes and a regression branch using a robust Smooth L1 or IoU loss on box offsets, combined as a weighted sum.
Deploying real-time detection on edge devices
WHAT IT TESTS: end-to-end edge deployment reasoning. OUTLINE: pick an efficient one-stage detector, train with augmentation, then quantize, prune, and compile to a hardware-accelerated runtime, measuring latency and accuracy tradeoffs.
Focal Loss and class imbalance in detectors
WHAT IT TESTS: handling extreme class imbalance. OUTLINE: focal loss multiplies cross-entropy by a (1-p)^gamma factor that down-weights easy, well-classified examples so the vast easy background does not swamp the loss.
Mean Average Precision in object detection
WHAT IT TESTS: the headline detection metric. OUTLINE: AP is the area under the precision-recall curve per class; mAP averages AP over classes, and COCO also averages over IoU thresholds. RED FLAG: confusing mAP with plain accuracy or ignoring the PR curve.
Intersection over Union for detection
WHAT IT TESTS: the core overlap metric. OUTLINE: IoU is the area of overlap divided by the area of union of predicted and ground-truth boxes; a threshold decides true positives.
Image classification versus object detection
WHAT IT TESTS: basic task definitions. OUTLINE: classification assigns one label to the whole image; detection localizes and labels multiple objects with bounding boxes and class scores.
Translation equivariance versus invariance in CNNs
WHAT IT TESTS: precise reasoning about CNN symmetries. OUTLINE: convolution is equivariant, shifting input shifts feature maps; invariance comes only from pooling and global aggregation. Strict invariance is partial and broken by strided sampling.
Depthwise separable convolution cost savings
WHAT IT TESTS: efficient convolution factorization. OUTLINE: separable conv splits standard conv into per-channel spatial filtering plus a 1x1 pointwise mix, cutting cost by roughly 1/N plus 1/k².
Adapting a classification CNN for segmentation
WHAT IT TESTS: turning a classifier into a dense predictor. OUTLINE: replace the dense head with conv layers, upsample via transposed convolutions, and fuse encoder skip connections to recover spatial detail lost to downsampling.
Uses of the 1x1 convolution
WHAT IT TESTS: channel-wise operations and efficient design. OUTLINE: a 1x1 conv is a per-pixel linear combination across channels; it reshapes channel depth cheaply and adds nonlinearity. Uses: dimensionality reduction in bottlenecks and channel mixing.
Receptive fields in convolutional networks
WHAT IT TESTS: how spatial context accumulates in CNNs. OUTLINE: receptive field is the input region affecting a neuron; it grows with depth, larger kernels, and stride. It matters for capturing context in detection and segmentation.
ResNet residual blocks and the degradation problem
WHAT IT TESTS: why skip connections enable very deep nets. OUTLINE: a residual block learns F(x) and adds the identity input x, so layers fit a residual; this eases gradient flow and solves the degradation problem.