All bites
The whole library, newest first. Filter by what you are here for, or pick a topic if you already know.
4330 bites
Page 216
The analysis phase: tokenizers and token filters
Analysis turns raw text into index terms via a tokenizer that splits text into tokens then token filters that transform them, like lowercasing or stemming.
TSM-Tree vs LSM-Tree storage engines
Both buffer writes in memory and flush sorted immutable files, but TSM organizes by series and time with columnar, heavily compressed blocks tuned for ordered appends and range scans.
Tuning HNSW for recall vs latency
ANN trades exactness for speed, and HNSW knobs M and efConstruction shape graph quality while efSearch trades query latency for recall at runtime.
Zero-downtime schema migration on a hot table
Expand-migrate-contract phases, dual-write and backfill, decouple deploys from migrations.
SQL isolation levels and the anomalies they prevent
Read Uncommitted allows dirty reads; Read Committed blocks them; Repeatable Read blocks non-repeatable reads; Serializable blocks phantoms.
When a graph database beats relational or document stores
Deeply connected data, variable-depth traversals, fraud or recommendation paths, index-free adjacency.
Choosing a good shard key and avoiding hot spots
High cardinality, even write distribution, query alignment; monotonic keys send all writes to one shard.
Synchronous vs asynchronous replication trade-offs
Sync waits for replica ack giving zero data loss but higher latency; async acks immediately, faster but risks losing recent writes on failover.
The N+1 query problem and how to fix it
One query for a list plus one per item for its relation, fix with eager loading or a batched join.
Tuning a database connection pool
Max size, min idle, connection and max-lifetime timeouts; size from cores and latency, not guesswork.
Upgrading a stateful Flink job without losing state
Take a savepoint, stop with drain, deploy new jar, restore from savepoint with matching operator UIDs.
Denormalization: trading write cost for read speed
Duplicate or precompute data to avoid joins, accept harder writes and consistency risk, justify by read-heavy access.
Predicate pushdown and why it speeds queries
Apply WHERE conditions at the scan or remote source, prune partitions and rows early, shrink data movement.
The buffer pool's role in database IO
Caches pages, serves reads from RAM, buffers dirty writes flushed later, uses eviction like LRU.
Vectorized query execution and its speedups
Process column batches per operator call, amortize per-tuple overhead, use cache locality and SIMD.
Sessionizing clickstream events into sessions
Order events per user, split on inactivity gap, assign session ids, pick event or session grain.
Diagnosing database latency layer by layer
Split total time into pool-wait, query execution, and ORM-generated query patterns; use metrics at each layer.
Choosing a time-series database for metrics
High-ingest timestamped writes, time-window queries, retention and downsampling, time-optimized compression.
Phantom reads and how serializable prevents them
New rows matching a predicate appear between reads; classic Repeatable Read locks existing rows not ranges; Serializable uses range or predicate locks.
The Volcano iterator model of query execution
Each operator exposes open/next/close, parents pull tuples from children, uniform composable interface, pipelined low memory.