tezvyn:

Computer Vision

Image/video models, diffusion, OCR, multimodal

301 bites

More in Computer Vision — page 3

Computer Vision80 sec read

Compare Gray World and White Patch white balance.

WHAT IT TESTS: classic color constancy methods. OUTLINE: Gray World assumes average scene color is gray, White Patch assumes the brightest pixel is white, both fail on dominant colors or clipping; learning predicts illuminant from data.

Computer Vision73 sec read

How does smartphone Portrait Mode produce bokeh?

WHAT IT TESTS: depth estimation plus synthetic rendering. OUTLINE: estimate per-pixel depth via dual-pixel or stereo or learning, segment the subject, then apply depth-dependent blur.

Computer Vision77 sec read

Outline the classic image stitching pipeline.

WHAT IT TESTS: feature-based image stitching. OUTLINE: detect and match features like SIFT, estimate a homography with RANSAC, warp and blend with multiband or feathering.

Computer Vision77 sec read

How do you build an HDR image from bracketed exposures?

WHAT IT TESTS: HDR imaging pipeline basics. OUTLINE: align frames, recover the camera response function, merge to a linear radiance map weighted by exposure, then tone map for display.

Computer Vision79 sec read

How do BYOL and Barlow Twins avoid representation collapse?

WHAT IT TESTS: self-supervised learning and collapse avoidance. OUTLINE: collapse is embeddings shrinking to a constant or low-rank subspace; BYOL uses predictor plus momentum target plus stop-gradient, Barlow Twins decorrelates feature dimensions.

Computer Vision82 sec read

How does MAML's inner and outer loop work?

WHAT IT TESTS: meta-learning and bi-level optimization. OUTLINE: inner loop does task-specific gradient steps from shared init, outer loop updates the init for fast adaptability via second-order gradients.

Computer Vision2 min read

Prototypical Networks for few-shot classification

WHAT IT TESTS: metric-based few-shot learning. OUTLINE: an encoder embeds support examples, each class prototype is the mean embedding of its support examples, and a query is classified by nearest prototype using a distance like Euclidean via softmax.

Computer Vision2 min read

Contrastive self-supervised learning with SimCLR

WHAT IT TESTS: contrastive learning mechanics. OUTLINE: two augmentations of one image form a positive pair, other images in the batch are negatives; an encoder plus projection head and the NT-Xent loss pull positives together and push negatives apart.

Computer Vision2 min read

Leveraging unlabeled data with limited labels

WHAT IT TESTS: semi-supervised and self-supervised strategy. OUTLINE: pretrain a representation on the million unlabeled images via self-supervision, then fine-tune on the 1,000 labels; or use pseudo-labeling and consistency regularization.

Computer Vision2 min read

Transfer learning from ResNet50 on small data

WHAT IT TESTS: applying transfer learning. OUTLINE: replace the final classification head with one sized to your classes, freeze the pretrained convolutional backbone as a feature extractor, train the new head, then optionally fine-tune top blocks at a low…

Computer Vision2 min read

Formulating a multi-step robot manipulation task

WHAT IT TESTS: end-to-end robot RL formulation. OUTLINE: perception detects and localizes the mug, action space spans navigation and manipulation, and a reward shaped over subgoals (reach, grasp, transport, place) with sparse final success guides learning.

Computer Vision2 min read

NeRF limitations and advances for robotics

WHAT IT TESTS: practical limits of NeRF. OUTLINE: original NeRF is slow to train and render, per-scene, static, and needs many calibrated views; address speed with explicit grids or Gaussian splatting, dynamics with time-conditioned fields, and scale with…

Computer Vision2 min read

Camera intrinsics, extrinsics, and the essential matrix

WHAT IT TESTS: epipolar geometry fundamentals. OUTLINE: intrinsics map camera coords to pixels, extrinsics are camera pose in the world; the essential matrix relates normalized points across two views, encoding relative rotation and translation up to scale…

Computer Vision2 min read

Designing a baseline Visual Question Answering model

WHAT IT TESTS: multimodal baseline design. OUTLINE: encode the image with a CNN, encode the question with an RNN or embedding, fuse the two vectors, and classify over a fixed answer vocabulary.

Computer Vision2 min read

Semantic, instance, and panoptic segmentation

WHAT IT TESTS: distinguishing segmentation paradigms. OUTLINE: semantic labels every pixel by class without separating objects; instance separates individual objects but may skip background; panoptic unifies both, labeling stuff and distinct thing instances.

Computer Vision2 min read

Designing a high-resolution photorealistic face generator

WHAT IT TESTS: system design for high-res faces. OUTLINE: weigh StyleGAN's fast, controllable style-based synthesis against diffusion's diversity and stable training; handle scale via progressive or multi-resolution synthesis; protect diversity to avoid mode…

Computer Vision2 min read

DDPM versus DDIM sampling trade-offs

WHAT IT TESTS: stochastic versus deterministic sampling. OUTLINE: DDPM is a stochastic Markov chain needing many steps; DDIM is a non-Markovian, deterministic sampler that skips steps for far faster inference and reproducible, invertible latents, trading a…

Computer Vision2 min read

Classifier-free guidance in diffusion models

WHAT IT TESTS: how guidance improves conditioning. OUTLINE: train one model jointly on conditional and dropped-condition inputs; at inference extrapolate from unconditional toward conditional prediction via a guidance scale, sharpening prompt adherence…

Computer Vision2 min read

Unpaired image translation with CycleGAN

WHAT IT TESTS: unpaired translation design. OUTLINE: CycleGAN uses two generators and two discriminators with a cycle-consistency loss that forces translating to the other domain and back to reconstruct the input, removing the need for paired data.

Computer Vision2 min read

How text prompts guide Stable Diffusion

WHAT IT TESTS: the text-conditioning pipeline. OUTLINE: a frozen text encoder turns the prompt into token embeddings, which feed the U-Net via cross-attention at each denoising step so the prompt steers generation; classifier-free guidance amplifies the…