Why Reliability and Observability Will Win the AI Era | Daniil Mazepin

The engineering leader on why reliability, observability and honest measurement, not features, will separate the winners of the AI era.As enterprise software races to bolt AI onto everything, one engineering leader cares more about what happens after the demo. Daniil Mazepin, a London-based Senior Engineering Manager with just over fifteen years across gaming, big tech and fintech, has spent his career on the parts of software users never see: reliability, observability, and whether a new capability actually works at scale.

His work sits at the intersection of large-scale distributed systems, the reliability and observability of AI, and the leadership needed to run both, three forces increasingly decisive for the next generation of software.

“Everyone can ship a clever feature now,” he tells Read. “The teams that win are the ones who can still tell you, six months later, whether it is reliable, what it costs, and whether it moved anything that matters.”

A consistent thread runs through Mazepin’s career he has led engineering where a single failure is felt by millions. At gaming giant Playtech he led a 30-strong mobile team, and a mobile redesign he built himself cut on-device storage by roughly 99 per cent, unlocking new markets. At Meta he worked on the global rollout of Shops and rolled out a reliability framework across teams. At British fintech unicorn Teya he built the cross-border payment platform that drove a roughly tenfold rise in eligible transactions and opened it to small businesses across Europe. His work has quietly touched millions of users across Europe, Asia and the United States.

Mazepin is also an IEEE Senior Member, a grade held by less than 50,000 professionals worldwide, and his writing on system reliability and AI adoption has reached hundreds of thousands. He speaks at conferences across Europe and mentors rising engineers in Africa through GMI. So when he argues that reliability, not features, will decide the next wave, it is the assessment of someone trusted with systems that cannot fail. Below are the key takeaways from his expert read on the trends reshaping software.

Why most AI features are already commodities
Mazepin believes much of the industry is fixating on the wrong layer. The AI capabilities that dominated product roadmaps a year ago, he argues, are no longer differentiators. Conversational assistants, semantic search and automated pricing engines have become table stakes, available to anyone with an API key and a weekend.

“Features are being commoditised faster than most roadmaps can keep up,” he argues. “If a competitor can rebuild your headline feature in a weekend, that feature was never your advantage.”

Where, then, does durable advantage sit? His philosophy of scale is contrarian: the edge is not the feature but the layer beneath, whether the system is reliable, observable and trustworthy under load. A model that tests well but degrades silently in production, he notes, is worse than no feature at all: it erodes the one thing that is hard to rebuild, user trust.

The dashboards that stopped working
Another trend Mazepin flags is less comfortable: the tools engineers use to know whether software is healthy were built for a world AI has quietly broken.

As Mazepin points out, traditional monitoring assumes deterministic systems, the same input yielding the same output, so a dashboard can flag when something deviates. AI systems do not behave that way. “A non-deterministic system can be green on every metric you have and still be quietly wrong,” he says. “Most teams are watching dashboards that no longer describe their software.”

The answer, he argues, is not to abandon observability but to rebuild it for probabilistic systems: tracing across chains of model calls, evaluation baked into production rather than run once before launch, and metrics that capture quality and drift, not just latency and errors. Much of the discipline transfers from large-scale engineering, he says, but the parts that assume determinism must be redesigned.

Measuring what AI agents actually deliver
A further shift he watches closely is agentic software: systems that do not just answer questions but carry out multi-step tasks within set limits. The promise is real, he believes, but so is the noise, and most teams are measuring the wrong thing.

“The trap is counting activity and calling it impact,” he says. “Lines of code, tickets closed, tasks automated, none of that tells you whether the business is better off.” His prescription is a discipline of three tests: volume, verifiability and value. Is the task frequent enough to automate; can you verify the output; and does it move a number anyone cares about?

Agentic coding: new tools, old rules?
The bigger risk, in Mazepin’s view, is subtler than any single tool: teams adopt AI while holding on to the habits it was meant to change. Handing a model a small, perfectly specified ticket that someone else wrote, he argues, is not real adoption; it is the old assembly line with a faster machine bolted on. The deeper shift he sees is one of mindset. As AI absorbs more of the question “am I building the thing right?”, the human edge moves towards “am I building the right thing?”, a matter of more ownership and product judgement, and of a higher form of quality discipline rather than less.

Review, too, is becoming an engineering problem in its own right. The teams adapting well, he observes, design for the new volume rather than fighting it, leaning on automation, test coverage and safety nets so that quality scales with the output instead of trailing behind it. “The tools are new, but the goal isn’t,” he says. “The habits have to change; the commitment to quality does not.”

The hidden social cost of speed
Mazepin puts as much emphasis on people as on technology, and here his most contrarian argument surfaces. AI is lifting the throughput of individual engineers, he observes, but he sees a second-order effect few leaders track: it is eroding the shared substrate that makes a team more than a collection of individuals.

“AI makes each engineer faster and, without anyone noticing, a little more alone,” he says. When a developer asks a model instead of a colleague, collaboration quietly falls, knowledge pools into silos, and the informal understanding that holds a system together thins out. “Your velocity dashboard will look excellent. It will not show you that your team stopped talking to each other.” The remedy, he argues, has to be deliberate, because the tools now pull the other way: leaders have to design for connection on purpose.

“That is why I actively promote a culture of curiosity, continuous learning and collaboration,” he says. “The tools make it easy to work alone, so keeping teams learning from one another has to be deliberate.”

Where the next wave will be won
According to Mazepin’s observation, where the industry is heading circles back to discipline. “The next wave will not belong to whoever ships AI first,” he says. “It will belong to the teams who can prove, with real numbers, that what they shipped actually works.” In a market intoxicated by launches, it is a sober bet, and an increasingly popular one.

Leave a Comment