Skip to content
tezvyn:

Computer Vision

Image/video models, diffusion, OCR, multimodal

72 bites

Test yourself: Top 30 intermediate Computer Vision interview questionsMultiple choice, with the correct answer and why it is correct on every question. Free, no sign-in.

Intermediate interview questions in Computer Vision, page 2

intermediate1 min read

How to improve coarse segmentation boundaries?

Skip connections and higher-resolution features, boundary-aware losses, and point-based or CRF refinement.

intermediate1 min read

Brightness constancy and small-motion assumptions

Brightness constancy says a point's intensity is invariant under motion; small motion lets you linearize via Taylor expansion.

intermediate1 min read

Kalman filter for bounding-box tracking

State, transition, measurement models, and process plus measurement noise; predict then correct each frame. State holds box position and velocity; measurement is the detected box.

intermediate1 min read

3D CNNs vs two-stream action recognition

3D CNNs learn spatiotemporal filters end to end but are heavy; two-stream splits RGB appearance and precomputed optical flow, strong but costly to compute flow.

intermediate1 min read

Re-identification in multi-object tracking

Re-ID matches an object to its prior id using appearance embeddings, not just position; store track features and match re-entering detections by embedding similarity.

intermediate1 min read

Inductive biases of ViT versus CNN

CNNs bake in locality and translation equivariance; a plain ViT has almost none beyond patch structure, so it must learn spatial relations from data, needing large datasets or strong pretraining.

intermediate2 min read

How Swin Transformer achieves linear attention

Swin computes attention within local non-overlapping windows of fixed size, making cost linear in patches, then shifts windows between layers so information crosses boundaries.

intermediate1 min read

Cross-attention for visual question answering

In cross-attention queries come from one modality and keys/values from the other, e.g. text queries attend over image features so the question selects relevant regions.

intermediate2 min read

Why ViTs need positional embeddings

Self-attention is permutation invariant so patch order is lost; positional embeddings restore spatial location. CNNs encode position implicitly via the fixed convolution grid.

intermediate2 min read

Mode collapse in GAN training

Generator produces few outputs ignoring data diversity, caused by chasing whatever fools the current discriminator; mitigate with minibatch discrimination, unrolled GANs, or Wasserstein loss.

intermediate2 min read

Evaluating generative models with FID versus IS

FID compares Inception feature distributions of real and fake images via Frechet distance between two Gaussians; it uses real data as reference and detects diversity issues, unlike IS which uses no real…

intermediate1 min read

Why U-Net skip connections matter for denoising

Skips carry high-resolution spatial detail from encoder to decoder, preserving fine structure lost in downsampling and easing gradient flow, which lets the model restore detail while removing…

intermediate2 min read

How text prompts guide Stable Diffusion

A frozen text encoder turns the prompt into token embeddings, which feed the U-Net via cross-attention at each denoising step so the prompt steers generation; classifier-free guidance amplifies the…

intermediate2 min read

Unpaired image translation with CycleGAN

CycleGAN uses two generators and two discriminators with a cycle-consistency loss that forces translating to the other domain and back to reconstruct the input, removing the need for paired data.

intermediate2 min read

Camera intrinsics, extrinsics, and the essential matrix

Intrinsics map camera coords to pixels, extrinsics are camera pose in the world; the essential matrix relates normalized points across two views, encoding relative rotation and translation up to scale…

intermediate2 min read

Contrastive self-supervised learning with SimCLR

Two augmentations of one image form a positive pair, other images in the batch are negatives; an encoder plus projection head and the NT-Xent loss pull positives together and push negatives apart.

intermediate2 min read

Prototypical Networks for few-shot classification

An encoder embeds support examples, each class prototype is the mean embedding of its support examples, and a query is classified by nearest prototype using a distance like Euclidean via softmax.

intermediate1 min read

How does smartphone Portrait Mode produce bokeh?

Estimate per-pixel depth via dual-pixel or stereo or learning, segment the subject, then apply depth-dependent blur.

intermediate1 min read

Compare Gray World and White Patch white balance.

Gray World assumes average scene color is gray, White Patch assumes the brightest pixel is white, both fail on dominant colors or clipping; learning predicts illuminant from data.

intermediate1 min read

How do you speed up a slow detection model?

Quantization, pruning, distillation, lighter backbones, and resolution or batching tweaks, each trading some accuracy or effort for speed.

We are hiring for this. Every open role lists the topics its interview covers, so you can prepare for the real thing rather than guessing.

See open roles