Intermediate interview questions in LLMs & Generative AI, page 4
Why RAG persists despite million-token context windows
Cost and latency scale with context, attention degrades in the middle, and RAG adds freshness, access control, and citations.
Self-attention and the Query, Key, Value matrices
Queries score against keys via scaled dot product, softmax yields weights, and those weight the values into the output.
How Transformers encode token position
Attention is permutation-invariant, so positional encodings (sinusoidal, learned, or rotary) are added or applied.
Cross-attention versus self-attention in encoder-decoder Transformers
Cross-attention draws Queries from the decoder and Keys/Values from the encoder, letting the decoder condition on the source.
Tokens and vocabulary-size tradeoffs
A token is a subword unit; larger vocab shortens sequences but bloats the embedding matrix, smaller vocab generalizes but lengthens sequences.
Fault-tolerant checkpointing for thousand-GPU pre-training
Checkpoint weights, optimizer state, RNG, and data position together; use asynchronous sharded writes and automated detect-restart-resume.
Prompt engineering to curb extraction hallucinations
Ground strictly in source, allow null for missing fields, enforce a schema, and use few-shot examples; acknowledge prompting cannot fully eliminate it.
Self-consistency over chain-of-thought
Sample multiple CoT paths at nonzero temperature and majority-vote the final answer; cost scales with the number of samples.
Why chain-of-thought helps large models but not small ones
Small models lack reliable multi-step reasoning, so CoT just adds error-prone steps; adapt by using few-shot/fine-tuning or distillation for small tiers.
Handling a 401 error in an LLM agent's tool call
Catch the tool error, return a structured observation to the LLM, and distinguish recoverable retries from terminal failures needing re-plan or escalation.
CLIP's contrastive objective and zero-shot classification
Train image and text encoders to align matched pairs and repel mismatched ones in a shared space; classify zero-shot by comparing an image to text prompts of class names.
Evaluating faithfulness and compositionality in multimodal models
Use targeted probes with hard negatives, attribute-relation binding tests, and structured grounding checks; note BLEU rewards surface overlap not correctness.
Red teaming a generative model
Deliberately probe for harmful outputs across categories, document jailbreaks, and automate with adversarial prompt generators plus classifier-based judging.
Preventing PII in LLM outputs: curation, fine-tuning, or guardrails
Favor post-processing guardrails as the enforceable last line, backed by data curation; note each layer's tradeoffs and that defense-in-depth is best.
Constitutional AI versus standard RLHF
A written principle set guides self-critique and revision, plus AI feedback (RLAIF) replaces human preference labels.
Penalizing sycophancy in a reward model
Sycophancy is reward proxy gaming where agreeableness substitutes for correctness; counter it with truth-anchored labels, perturbed-premise pairs, and consistency checks.
The alignment tax and capability trade-offs
Alignment tax is capability lost from safety tuning, measured as benchmark or task-success deltas before and after; a product decision weighs over-refusal against harm risk.
Scalable oversight of superhuman models
Humans cannot judge outputs beyond their expertise, so feedback degrades; techniques like AI debate or recursive reward modeling decompose judgment.
Designing an LLM red-teaming framework
Taxonomy of harms, automated adversarial prompt generation via attacker models and mutation, a classifier to triage outputs, and severity-by-likelihood prioritization.
AWS Bedrock versus a direct provider API
Bedrock unifies many models with IAM, VPC, and cloud integration; a direct provider API gives earliest models, full feature parity, and simpler vendor terms.
We are hiring for this. Every open role lists the topics its interview covers, so you can prepare for the real thing rather than guessing.
See open roles