A Field Guide to Log Analysis
Monitoring Alerts: Periodic jobs should be safe to run twice, because they will be. Monitoring Alerts: You rarely need a new component to fix a boundary problem. Monitoring Alerts: The signal you want is often already logged, just not aggregated.
Backup Strategy: The interesting number is not the average, it is the 99th percentile. Backup Strategy: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Backup Strategy: Every abstraction you add is a place where behaviour can differ from intent.
Log Analysis: If the rollback plan needs a meeting, it is not a rollback plan. Log Analysis: Small pages that stay small are easier to keep fast than large ones made fast. Log Analysis: Write the invariant down; otherwise it lives only in someone's memory.
Search Indexing: Configurations should be reviewable in a diff, not only in a console. Search Indexing: The best time to add an index is before the table gets large. Search Indexing: Failures are usually correlated, so plan for the shared dependency.
For access control, the constraint matters more than the feature list. Periodic jobs should be safe to run twice, because they will be. Teams working on access control usually discover this the hard way. You rarely need a new component to fix a boundary problem. The signal you want is often already logged, just not aggregated. This is most visible in access control.
API Design: The first thing to settle is the failure mode, not the happy path. API Design: Measurements taken once are anecdotes; you need a baseline that repeats. API Design: Costs usually concentrate in a small number of operations, so find those first.
Cloud Infrastructure: Periodic jobs should be safe to run twice, because they will be. You rarely need a new component to fix a boundary problem. That applies to cloud infrastructure as well. In practice, cloud infrastructure behaves differently: The signal you want is often already logged, just not aggregated.
In practice, log analysis behaves differently: The first thing to settle is the failure mode, not the happy path. Measurements taken once are anecdotes; you need a baseline that repeats. The same reasoning holds for log analysis. For log analysis, the constraint matters more than the feature list. Costs usually concentrate in a small number of operations, so find those first.
A boundary is a limit a person sets around their own body, time, privacy or emotional wellbeing. In a relationship, it might concern which kinds of physical contact feel welcome, whether a person wants to use a barrier method during sex, how personal information is shared, or when they need time alone. Boundaries can be broad, but clear examples are easier to understand and respect.
Log Analysis: You can often replace a coordination problem with an idempotency key. Log Analysis: Anything that grows without a bound will eventually hit one. Log Analysis: Documentation that is not tested tends to describe the previous version.
Cost Controls: If a metric has no owner, it will drift until it causes an incident. Cost Controls: The cheapest optimisation is usually removing work nobody asked for. Cost Controls: Aggregating at write time trades flexibility for predictable read cost.
Consent requires the ability to make and communicate a choice. Someone who is asleep or unconscious cannot agree at that time. Alcohol or other drugs can affect judgment and communication, but the legal rules for assessing capacity vary. The relevant question is not simply whether someone has consumed a substance; it is whether they can understand the choice and make it freely. If that is unclear, do not proceed.
In practice, observability behaves differently: Configurations should be reviewable in a diff, not only in a console. The best time to add an index is before the table gets large. The same reasoning holds for observability. For observability, the constraint matters more than the feature list. Failures are usually correlated, so plan for the shared dependency.
Access Control: The interesting number is not the average, it is the 99th percentile. Access Control: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Access Control: Every abstraction you add is a place where behaviour can differ from intent.
Edge Caching: If a metric has no owner, it will drift until it causes an incident. Edge Caching: The cheapest optimisation is usually removing work nobody asked for. Edge Caching: Aggregating at write time trades flexibility for predictable read cost.
In practice, schema migration behaves differently: If a metric has no owner, it will drift until it causes an incident. The cheapest optimisation is usually removing work nobody asked for. The same reasoning holds for schema migration. For schema migration, the constraint matters more than the feature list. Aggregating at write time trades flexibility for predictable read cost.
In practice, schema markup behaves differently: A queue smooths spikes but also hides how far behind you are. Retries without jitter turn a small outage into a large one. The same reasoning holds for schema markup. For schema markup, the constraint matters more than the feature list. Separating the reads from the writes buys room to change either side.
Schema Migration: Periodic jobs should be safe to run twice, because they will be. Schema Migration: You rarely need a new component to fix a boundary problem. Schema Migration: The signal you want is often already logged, just not aggregated.
For content delivery, the constraint matters more than the feature list. A queue smooths spikes but also hides how far behind you are. Teams working on content delivery usually discover this the hard way. Retries without jitter turn a small outage into a large one. Separating the reads from the writes buys room to change either side. This is most visible in content delivery.
API Design: You can often replace a coordination problem with an idempotency key. API Design: Anything that grows without a bound will eventually hit one. API Design: Documentation that is not tested tends to describe the previous version.
Periodic jobs should be safe to run twice, because they will be. This is most visible in data pipelines. Consider data pipelines specifically. You rarely need a new component to fix a boundary problem. Data Pipelines: The signal you want is often already logged, just not aggregated.
A queue smooths spikes but also hides how far behind you are. This is most visible in api design. Consider api design specifically. Retries without jitter turn a small outage into a large one. API Design: Separating the reads from the writes buys room to change either side.
Teams working on storage tiers usually discover this the hard way. The interesting number is not the average, it is the 99th percentile. Adding a cache in front of a slow query is a fix; fixing the query is a cure. This is most visible in storage tiers. Consider storage tiers specifically. Every abstraction you add is a place where behaviour can differ from intent.
Monitoring Alerts: If the rollback plan needs a meeting, it is not a rollback plan. Monitoring Alerts: Small pages that stay small are easier to keep fast than large ones made fast. Monitoring Alerts: Write the invariant down; otherwise it lives only in someone's memory.