Skip to main content

Akmatori Blog

Blog about SRE, DevOps, Linux and Networks

Prometheus Full Disk Crash Fix - blog post cover image
10.08.2026Prometheus Full Disk Crash Fix

Prometheus v3.13.2 fixes a PromQL SIGBUS crash path that could appear when the data disk is full. For SRE teams, the lesson is simple: your monitoring system needs the same disk-failure discipline as the services it watches.

Continue reading
MySQL SKIP LOCKED at Scale - blog post cover image
09.08.2026MySQL SKIP LOCKED at Scale

Shopify's inventory reservation rewrite is a useful reminder for SRE teams: the simplest reliable system is often the one that removes a moving part. MySQL 8, careful lock design, and connection visibility let Shopify replace ...

Continue reading
AI Scraper Overload Is an SRE Problem - blog post cover image
08.08.2026AI Scraper Overload Is an SRE Problem

AI crawlers are no longer a background annoyance. The Gentoo Bugzilla incident is a useful reminder that public engineering systems need crawler budgets, static exports, and overload runbooks before automated clients find them...

Continue reading
  • ...

Automate incident response and prevent on-call burnout with AI-driven agents!