War Stories

Real production incidents, anonymized. No client names, no screenshots, no repo links. Just the patterns that break systems — and the fixes that save them.

SEV1 Response Automotive Kafka / Debezium

The Kafka Consumer Group Bug That Looked Like a Network Problem

The Problem

A platform processing 50K+ events/minute experienced cascading CDC failures. Debezium connectors stalled, consumer lag spiked to 6+ hours, and downstream services timed out. The on-call team had been chasing "network issues" for 6 hours.

What I Did

The Outcome

Recovery within 4 hours. Consumer lag normalized. System stable under 2x peak load. Team retained a 12-month advisory retainer.

Migration Assessment Enterprise WildFly / Kubernetes

From WildFly to Kubernetes: A Migration Assessment Framework

The Problem

A mid-size enterprise spent 3 months on a "lift-and-shift" Docker approach for their JBoss suite. Result: memory leaks, 8-minute startup times, and failed health checks.

What I Did

The Outcome

First two services migrated in 6 weeks. Memory usage dropped 40%. Startup time: 8 minutes → 45 seconds.

MLOps / CI-CD Healthtech AWS / Terraform

The ML Pipeline That Worked Locally But Broke in Production

The Problem

A healthtech startup's inference pipeline worked in Jupyter but failed silently in AWS ECS. Model artifacts weren't versioned. Feature stores were out of sync. No rollback strategy.

What I Did

The Outcome

Inference failures dropped from ~15% to under 0.5%. Model rollbacks became one-click. Data science team could deploy without DevOps.

Facing something similar?

Start a Conversation