top of page


When Your System Lies in Complete Sentences: Observing LLMs in Production
Picture this scenario. Your AI-powered feature is running. Every metric you watch looks clean. Error rate: flat. P99 latency: within SLO. HTTP 200s across the board. Your on-call engineer has nothing to page about. But for the last three days, the model has been producing subtly wrong output. Not wrong in a way that crashes anything. Not wrong in a way that fires a single alert. The responses are fluent, confident, and structurally perfect. They just happen to be incorrect in
Chandra Sekar Reddy
May 248 min read


Observability Is Not a Dashboard: What a Decade in SRE Taught Me About Modern Systems
Over the last ten years in Site Reliability Engineering, I have watched one word rise from technical obscurity to boardroom vocabulary: observability . It is now everywhere. Vendors lead with it. Engineering leaders budget for it. Teams build entire platforms around it. Every organization claims to be investing in it. And yet, in far too many environments, when a major production incident happens, the same confusion still surfaces within minutes: What changed? Where did it st
Chandra Sekar Reddy
Mar 89 min read
bottom of page