About
I'm Matthew Herzog, an engineering leader based in Seattle. I've spent most of my career in site reliability engineering: keeping large distributed systems running, and building the teams that do it.
Most recently I was a Senior SRE Engineering Leader at NVIDIA, on Hardware Infrastructure. I set up its Operational Excellence program, which covers more than 500 engineers: standards for incident management, maintenance and retrospectives, plus regular operational health reviews. I also built an LLM- and MCP-based incident analysis tool that pulls from Rootly, Jira and Google Drive to correlate incidents, draft retrospective summaries and track follow-up work.
Before that I spent almost 19 years at Google. I started on Windows infrastructure, moved to SRE for Google's source, build and test systems, and then managed SRE teams for internal business systems, for Google Cloud's managed storage products (Memorystore, Filestore, NetApp Volumes), and for the Cloud Storage metadata and transfer infrastructure behind services like YouTube, Drive and Gmail.
Earlier, I was a systems engineer at Premera Blue Cross and an engineer at US West. I have a BS in Electrical Engineering from New Mexico Tech and am an NVIDIA-Certified Associate in AI Infrastructure and Operations.
This blog is where I write about AI: agentic tools, automating operations work, and what it's like to build things with AI in practice. The site itself is one example. I built it with Claude, and this post explains how.
You can find me on LinkedIn.