Intermediate interview questions in Computer Vision, page 3
How do you train and evaluate on imbalanced defect data?
Resampling, class weighting, focal loss, and anomaly framing for training; evaluate with precision, recall, PR-AUC, and F-beta, not accuracy.
CPU vs GPU vs Edge TPU for inference.
CPU is flexible but slow, GPU offers massive parallelism at high power, Edge TPU gives efficient low-power int8 inference but is constrained; choose by latency, power, cost, and model fit.
How is an HDR radiance map constructed from exposures?
Recover the inverse camera response function from corresponding pixels, linearize each exposure to radiance, then merge with confidence weights into a floating-point radiance map.
Feature detector vs feature descriptor.
A detector finds where interesting points are, a descriptor encodes the local appearance around each so points can be matched.
Why learn detection and description jointly like SuperPoint?
A shared backbone jointly optimizes detection and description for matching, sharing computation and learning data-driven robustness instead of hand-crafted heuristics.
What is Bundle Adjustment and why is it tractable?
Jointly refine 3D points and camera poses by minimizing reprojection error, expensive due to many coupled parameters; sparsity of the Jacobian and the Schur complement make it tractable.
Improving small object detection
Raise input resolution and tile, use feature pyramids for high-res features, tune anchors and copy-paste augmentation.
Deploying segmentation on edge devices
Pick efficient architectures, apply INT8 quantization, distill from a large teacher.
Adapting a 2D CNN for video action recognition
Run the 2D CNN per frame, pool features over time, optionally add two-stream or 3D conv.
Core components of visual SLAM
Tracking estimates per-frame pose, mapping builds and refines the 3D map, loop closure detects revisits and corrects drift.
The data association problem in SLAM
Matching observations to landmarks, why wrong matches corrupt the map, robust techniques like RANSAC and descriptor matching.
Cross-attention in transformer VQA models
Text queries attend over image regions, learning alignment that grounds words to visual content.
Point cloud vs voxel grid vs NeRF
Point clouds are sparse and fast but unstructured, voxels are regular for collision checks but memory-heavy, NeRFs render photorealistically but are slow.
Zero-shot classification with CLIP
Encode image and label prompts into a shared space, compare via cosine similarity, pick the highest.
Contrastive learning vs masked image modeling
Contrastive aligns augmented views via instance discrimination; MAE reconstructs masked patches; they differ in augmentation and fine-tuning.
When a homography is a valid model
Homography holds for pure rotation or a planar scene; it fails with translation plus 3D parallax, where epipolar geometry applies.
Challenges deploying a model on edge hardware
Limited memory and compute cause latency, thermal and power limits, accuracy loss from compression, operator support gaps.
Designing a multi-object tracker
Detect per frame, predict motion with a filter, associate via IoU and appearance, manage track lifecycle, handle occlusion with re-ID.
One-stage vs two-stage detectors
One-stage predicts boxes directly for speed; two-stage proposes then refines for accuracy; focal loss narrows the gap.
Non-maximum suppression in detection
Detectors emit many overlapping boxes per object; NMS keeps the highest-scoring box and removes others above an IoU threshold.
We are hiring for this. Every open role lists the topics its interview covers, so you can prepare for the real thing rather than guessing.
See open roles