AI incident response your SREs can operate
Akmatori turns production alerts into governed response workflows: it gathers logs, metrics, traces, recent deploy context, and runbook knowledge, then proposes or executes the next safe action with human approval gates. Teams can self-host it, connect existing observability and on-call tools, and prove value on real incidents before expanding automation.
What Akmatori Does
The capabilities that ship in the open-source platform today
Multi-Source Alert Ingestion
Ingest webhooks from Alertmanager, PagerDuty, Grafana, Datadog, and Zabbix, plus Slack channel monitoring — no rip-and-replace of your tooling.
Agentic Investigation
Every incident runs a full agent session that gathers logs, metrics, and context, with subagents for runbook lookup and cross-incident memory recall.
Real Tools, Not Just Chat
Agents reach your infrastructure through SSH, Kubernetes, scripts, and API clients — reusable tools you scope per skill or cron job.
Approval-Gated Remediation
Risky actions become proposals you approve in the web UI or Slack before anything executes. You decide how much autonomy each agent gets.
Bring Any LLM
OpenAI, Anthropic, Google, OpenRouter, NVIDIA NIM, MiniMax, or any OpenAI-compatible on-prem endpoint (GLM, Kimi, Mistral, LLaMA). Swap providers anytime — no lock-in.
Scheduled Investigations
Cron jobs run recurring agent investigations with their own prompt and tool allowlist, posting results straight to a Slack channel.
Self-Hosted by Design
Deploy with Docker Compose or Helm. Your incident data never leaves your infrastructure — run fully on-prem with local models for air-gapped setups.
Web Dashboard + Slack
Manage incidents, skills, tools, and settings from a modern, mobile-responsive UI, and drive the same workflows from Slack.
