Autoconfig Independent coverage of news

Data Pipelines Explained Without the Jargon

By David Kim · · 1094 words
Data Pipelines Explained Without the Jargon

If a metric has no owner, it will drift until it causes an incident. This is most visible in observability. Consider observability specifically. The cheapest optimisation is usually removing work nobody asked for. Observability: Aggregating at write time trades flexibility for predictable read cost.

Release Process: Configurations should be reviewable in a diff, not only in a console. Release Process: The best time to add an index is before the table gets large. Release Process: Failures are usually correlated, so plan for the shared dependency.

Edge Caching: Serving static bytes is the cheapest thing you can do at the edge. Edge Caching: A schema is an interface; changing it is a migration, not an edit. Edge Caching: Track the denominator as carefully as the numerator.

Storage Tiers: The first thing to settle is the failure mode, not the happy path. Storage Tiers: Measurements taken once are anecdotes; you need a baseline that repeats. Storage Tiers: Costs usually concentrate in a small number of operations, so find those first.

Monitoring Alerts: Periodic jobs should be safe to run twice, because they will be. Monitoring Alerts: You rarely need a new component to fix a boundary problem. Monitoring Alerts: The signal you want is often already logged, just not aggregated.

Crawl Budget: Periodic jobs should be safe to run twice, because they will be. Crawl Budget: You rarely need a new component to fix a boundary problem. Crawl Budget: The signal you want is often already logged, just not aggregated.

Observability: Configurations should be reviewable in a diff, not only in a console. Observability: The best time to add an index is before the table gets large. Observability: Failures are usually correlated, so plan for the shared dependency.

Observability: The first thing to settle is the failure mode, not the happy path. Observability: Measurements taken once are anecdotes; you need a baseline that repeats. Observability: Costs usually concentrate in a small number of operations, so find those first.

For data pipelines, the constraint matters more than the feature list. If a metric has no owner, it will drift until it causes an incident. Teams working on data pipelines usually discover this the hard way. The cheapest optimisation is usually removing work nobody asked for. Aggregating at write time trades flexibility for predictable read cost. This is most visible in data pipelines.

Configurations should be reviewable in a diff, not only in a console. This is most visible in edge caching. Consider edge caching specifically. The best time to add an index is before the table gets large. Edge Caching: Failures are usually correlated, so plan for the shared dependency.

Release Process: Periodic jobs should be safe to run twice, because they will be. Release Process: You rarely need a new component to fix a boundary problem. Release Process: The signal you want is often already logged, just not aggregated.

Crawl Budget: Serving static bytes is the cheapest thing you can do at the edge. Crawl Budget: A schema is an interface; changing it is a migration, not an edit. Crawl Budget: Track the denominator as carefully as the numerator.

Boundaries may involve practical health decisions as well as personal comfort. If relevant, discuss contraception, barrier methods, STI testing, and what each person understands about risk before sexual activity. Be clear about what you will do if you cannot agree on a safety measure: for example, you may decide not to proceed. Neither partner should be expected to accept a risk they have not agreed to.

Cloud Infrastructure: The first thing to settle is the failure mode, not the happy path. Measurements taken once are anecdotes; you need a baseline that repeats. That applies to cloud infrastructure as well. In practice, cloud infrastructure behaves differently: Costs usually concentrate in a small number of operations, so find those first.

Teams working on api design usually discover this the hard way. A design that cannot be rolled back is a design that cannot be changed safely. Latency budgets are easier to defend when every hop has a stated ceiling. This is most visible in api design. Consider api design specifically. Caching helps only until the invalidation rules become the bottleneck.

Schema Migration: Serving static bytes is the cheapest thing you can do at the edge. Schema Migration: A schema is an interface; changing it is a migration, not an edit. Schema Migration: Track the denominator as carefully as the numerator.

Storage Tiers: Serving static bytes is the cheapest thing you can do at the edge. Storage Tiers: A schema is an interface; changing it is a migration, not an edit. Storage Tiers: Track the denominator as carefully as the numerator.

The first thing to settle is the failure mode, not the happy path. This is most visible in backup strategy. Consider backup strategy specifically. Measurements taken once are anecdotes; you need a baseline that repeats. Backup Strategy: Costs usually concentrate in a small number of operations, so find those first.

The interesting number is not the average, it is the 99th percentile. That applies to rate limiting as well. In practice, rate limiting behaves differently: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Every abstraction you add is a place where behaviour can differ from intent. The same reasoning holds for rate limiting.

Content Delivery: Periodic jobs should be safe to run twice, because they will be. Content Delivery: You rarely need a new component to fix a boundary problem. Content Delivery: The signal you want is often already logged, just not aggregated.

Log Analysis: Serving static bytes is the cheapest thing you can do at the edge. Log Analysis: A schema is an interface; changing it is a migration, not an edit. Log Analysis: Track the denominator as carefully as the numerator.

Edge Caching: If a metric has no owner, it will drift until it causes an incident. Edge Caching: The cheapest optimisation is usually removing work nobody asked for. Edge Caching: Aggregating at write time trades flexibility for predictable read cost.

Backup Strategy: If the rollback plan needs a meeting, it is not a rollback plan. Backup Strategy: Small pages that stay small are easier to keep fast than large ones made fast. Backup Strategy: Write the invariant down; otherwise it lives only in someone's memory.

Storage Tiers: A queue smooths spikes but also hides how far behind you are. Storage Tiers: Retries without jitter turn a small outage into a large one. Storage Tiers: Separating the reads from the writes buys room to change either side.

Related reading