Time-Series Databases
Time-Series Databases (TSDBs) are built to handle sequences of data points indexed in time order. From server telemetry (CPU utilization, memory, latency) to financial market ticks and IoT sensor metrics, time-series data is append-only, high-volume, and time-stamped.
Core Architecture & Optimizations
1. Sequential Append-Only Writes
Traditional relational databases spend CPU/disk resources performing random B-Tree index updates. TSDBs use LSM-Trees or Columnar MergeTrees (e.g. ClickHouse, InfluxDB) that turn random writes into fast sequential disk append operations.
2. Time-Based Partitioning (Hypertables)
Data is automatically sliced into immutable time chunks (e.g. 2-hour or 1-day blocks). Expired data can be purged instantly by dropping the partition block file on disk (DROP PARTITION) rather than executing expensive SQL DELETE queries that cause index bloat.
3. Specialized Compression
- Delta-of-Delta: Timestamp differences ($\Delta_1 - \Delta_0$) for fixed-interval metrics (e.g., every 10 seconds) compress to 1 bit per sample.
- Gorilla Float Compression: Compresses 64-bit floating point metric values using XOR bitwise differences down to ~1.37 bytes.
graph TD
A["Incoming Metric Stream<br/>50,000 writes/sec"] --> B["In-Memory Memtable<br/>Gorilla + Delta-of-Delta Compression"]
B --> C["Immutable Chunk File<br/>2-Hour Time Partition"]
C --> D["Background Compaction<br/>Merge & Downsample"]
The High-Cardinality Bottleneck
Cardinality is the number of unique time series generated by tag combinations:
Cardinality = Unique Services * Hosts * Endpoints * Status Codes
If application developers accidentally inject dynamic variables like user_id or uuid into metric labels, cardinality explodes into millions of streams, overwhelming inverted indexes in RAM and causing OOM crashes.
Key Takeaways
- TSDBs handle append-only, time-ordered data with 10β100x better compression than RDBMS.
- Time-partitioning enables instant retention drops via partition file deletion.
- Guard against High Cardinality by keeping metric tags strictly limited to low-cardinality enums.
In the next section, we explore Search Engines and Full-Text Search.
β‘ Revision Mode Active Condensed for 5-minute review. Focused on 30-second pitches, mental anchors, and high-frequency interview traps.
β‘ 30-Second Elevator Pitch
βTime-Series Databases (TSDBs like Prometheus, InfluxDB, TimescaleDB, ClickHouse) are optimized for high-volume, append-only, time-stamped metric streams (IoT sensors, server telemetry, financial ticks). Using LSM Trees (or Columnar MergeTrees), Time-Partitioned Sharding, and specialized encoding algorithms (Delta-of-Delta, Gorilla Float Compression), TSDBs achieve 10 to 100x compression and sub-millisecond range aggregation over billions of data points.β
π§ Under The Hood & Mental Model
- LSM Trees & Time-Partitioned Hypertables:
- Writes bypass B-Tree index locking by appending to in-memory Memtables flushed to disk as immutable, time-partitioned chunks (e.g. 2-hour blocks). Old partitions can be dropped instantly (
DROP PARTITION) without expensive row-by-row deletions.
- Compression Algorithms:
- Delta-of-Delta Encoding: Stores variations in timestamp intervals (
T_n - T_(n-1)). Constant 10s intervals compress down to 1 bit.
- Gorilla Floating-Point Compression: Stores XOR differences between consecutive float values. Sparse XOR bits compress values from 8 bytes down to 1.37 bytes.
- Run-Length Encoding (RLE): Compresses repeated static gauge values (e.g.
80% CPU over 100 samples).
- The High-Cardinality Explosion:
- Cardinality = Total unique label/tag key-value combinations (
Region * Service * Host * Endpoint).
- Exponential tag proliferation inflates in-memory inverted index trees, causing RAM exhaustion and crash-loops.
π» High-Yield Code & Gotcha
-- TimescaleDB Continuous Aggregate (Automated Downsampling Rollup)
CREATE MATERIALIZED VIEW metrics_hourly
WITH (timescaledb.continuous) AS
SELECT
time_bucket('1 hour', time) AS hour_bucket,
host_id,
avg(cpu_utilization) AS avg_cpu,
max(cpu_utilization) AS max_cpu
FROM metrics
GROUP BY hour_bucket, host_id;
-- Automated Retention Policy (Drop raw data older than 7 days)
SELECT add_retention_policy('metrics', INTERVAL '7 days');
β οΈ GOTCHA: Inserting Unbounded High-Cardinality Variables into Metric Labels
# β Error: Anti-Pattern (Putting User IDs or UUIDs into Prometheus Labels)
http_requests_total{service="payment", user_id="usr_982348923"} 1 # 10 Million Users = 10 Million Inverted Index Nodes!
# β
Safe: Idiomatic Pattern (Use Bounded Metric Labels; Move High-Cardinality IDs to Logs/Traces)
http_requests_total{service="payment", status_code="500"} 1
βοΈ Trade-Offs & Comparison
| Characteristic | General Relational DB (Postgres) | Time-Series Database (InfluxDB / Timescale) |
|---|
| Write Model | Random $O(\log N)$ B-Tree updates | Sequential Append-Only LSM / MergeTree |
| Compression Ratio | 1.5β2x standard page | 10β100x (Gorilla / Delta-of-Delta) |
| Old Data Removal | Expensive DELETE WHERE timestamp < X | Instant partition file drop (DROP TABLE shard) |
| Primary Query Type | High-selectivity Point Lookup | Time-bucketed Range Aggregations (AVG, P95) |
| Cardinality Handling | Excellent (Indexed Primary Keys) | Vulnerable to RAM crash on Tag Explosions |
π My Experience (STAR Anchor)
βEngineered the metrics monitoring platform for a fleet of 20,000 IoT edge devices. Initial attempts to store 50,000 metrics/sec in standard PostgreSQL crashed DB disks due to index maintenance. By migrating to TimescaleDB with continuous aggregates and Gorilla compression, we reduced disk I/O by 90%, achieved a 14x storage reduction, and cut P99 dashboard latency from 8s to 12ms.β