Over the past few months, I've been speaking with Heads of Engineering, Engineering Managers, SREs, Platform Engineers, Security Engineering leaders, and architects while researching how engineering teams respond to major production incidents.
I originally assumed the biggest challenge would be observability.
Modern engineering organizations already have excellent tools for logs, metrics, traces, monitoring, and alerting.
But the conversations kept pointing somewhere else.
The recurring challenge wasn't seeing that something had gone wrong.
It was getting multiple teams to quickly reach the same understanding of what was happening so they could coordinate an effective response.
Different leaders described it in different ways:
That made me question whether we've been focusing on only part of the problem.
Observability helps teams detect and investigate.
But there still seems to be a gap between detecting an incident and converging on a coordinated response.
I'm starting to think that "operational convergence"—the ability for teams to rapidly establish shared understanding, ownership, and coordinated execution—may deserve to be treated as its own engineering capability rather than simply another feature of observability.
I'm still validating this idea, so I'm curious how others see it.
For those operating distributed systems:
Do you think operational convergence is already adequately solved by today's observability platforms, or is it a separate problem that engineering organizations still struggle with?