More in AI & ML — page 13
Core principles of a Neural Radiance Field
WHAT IT TESTS: implicit 3D representation and differentiable rendering. OUTLINE: an MLP maps a 3D point plus view direction to color and density; novel views render by casting rays, sampling points, querying the MLP, and volume-integrating along each ray.
Why ViTs need positional embeddings
WHAT IT TESTS: why order matters for attention but not convolution. OUTLINE: self-attention is permutation invariant so patch order is lost; positional embeddings restore spatial location. CNNs encode position implicitly via the fixed convolution grid.
Cross-attention for visual question answering
WHAT IT TESTS: fusing two modalities with attention. OUTLINE: in cross-attention queries come from one modality and keys/values from the other, e.g. text queries attend over image features so the question selects relevant regions.
How Swin Transformer achieves linear attention
WHAT IT TESTS: making attention scale to high resolution. OUTLINE: Swin computes attention within local non-overlapping windows of fixed size, making cost linear in patches, then shifts windows between layers so information crosses boundaries.
Inductive biases of ViT versus CNN
WHAT IT TESTS: how built-in priors affect data needs. OUTLINE: CNNs bake in locality and translation equivariance; a plain ViT has almost none beyond patch structure, so it must learn spatial relations from data, needing large datasets or strong pretraining.
Self-attention over image patches explained
WHAT IT TESTS: the QKV mechanism on patch tokens. OUTLINE: each patch projects to query, key, value; a patch's query is scored against all keys, softmax-normalized into weights, used to combine all values.
How ViT and CNN process an image differently
WHAT IT TESTS: the input pipelines of two paradigms. OUTLINE: a CNN slides local filters over the raw pixel grid; a ViT splits the image into patches, flattens and linearly embeds each into a token, adds positional embeddings, and feeds the sequence to…
Scene flow versus optical flow
WHAT IT TESTS: 2D versus 3D motion estimation. OUTLINE: optical flow is 2D pixel motion in the image plane; scene flow is the 3D motion field of points in space, needing depth via stereo, RGB-D, or LiDAR.
Self-supervised pretraining for video understanding
WHAT IT TESTS: learning from unlabeled video via pretext tasks. OUTLINE: define a label-free task like temporal order prediction or contrastive clip matching that forces temporal reasoning, then fine-tune on labeled action data.
Re-identification in multi-object tracking
WHAT IT TESTS: maintaining identity across gaps and occlusions. OUTLINE: Re-ID matches an object to its prior id using appearance embeddings, not just position; store track features and match re-entering detections by embedding similarity.
3D CNNs vs two-stream action recognition
WHAT IT TESTS: how architectures capture temporal motion. OUTLINE: 3D CNNs learn spatiotemporal filters end to end but are heavy; two-stream splits RGB appearance and precomputed optical flow, strong but costly to compute flow.
Kalman filter for bounding-box tracking
WHAT IT TESTS: predict-update recursion applied to tracking. OUTLINE: state, transition, measurement models, and process plus measurement noise; predict then correct each frame. State holds box position and velocity; measurement is the detected box.
Brightness constancy and small-motion assumptions
WHAT IT TESTS: foundations and limits of optical flow. OUTLINE: brightness constancy says a point's intensity is invariant under motion; small motion lets you linearize via Taylor expansion.
Design a tracking-by-detection tracker
WHAT IT TESTS: building tracking from detection plus data association. OUTLINE: detect per frame, then associate boxes across frames by IoU or appearance using Hungarian matching, maintaining track ids.
Sparse vs dense optical flow and Lucas-Kanade
WHAT IT TESTS: understanding motion estimation granularity. OUTLINE: sparse flow tracks selected feature points, dense flow computes a vector per pixel; Lucas-Kanade solves brightness constancy in a local window assuming constant motion.
Adapting ViT for dense semantic segmentation
WHAT IT TESTS: turning a classification ViT into a dense predictor. OUTLINE: reassemble patch tokens into a 2D feature map, add a decoder, and handle low resolution plus quadratic attention cost.
Explain panoptic segmentation and Panoptic Quality
WHAT IT TESTS: unifying semantic and instance segmentation plus its metric. OUTLINE: panoptic assigns every pixel a class and instance id over things and stuff; PQ factors into SQ, average IoU of matches, times RQ, an F1 over matched segments.
How to improve coarse segmentation boundaries?
WHAT IT TESTS: practical debugging of low-resolution mask edges. OUTLINE: skip connections and higher-resolution features, boundary-aware losses, and point-based or CRF refinement.
How does Mask R-CNN do instance segmentation?
WHAT IT TESTS: understanding of two-stage detectors and per-instance masks. OUTLINE: Faster R-CNN backbone plus RPN, then RoIAlign and a parallel mask head predicting per-class binary masks. RED FLAG: claiming masks are shared or that RoIPool is used.
U-Net architecture and its skip connections
WHAT IT TESTS: encoder-decoder design for segmentation. OUTLINE: U-Net has a contracting encoder, an expanding decoder, and skip connections that concatenate matching-resolution encoder features into the decoder to recover spatial detail lost in downsampling.