A production alert is not a request for an agent to fix everything it can reach. Start with evidence. An agent can sort logs, distinguish a hypothesis from an observed failure, and prepare a mitigation for review. But permissions and escalation rules need to be settled before the pager goes off.

These four entries in the AgentNDX skills directory cover distinct points in an incident workflow. They are not a substitute for observability, a runbook, or an incident commander.

What to look for

Scope and authority. Does the skill merely analyze evidence, or can it change production state? A triage assistant should not inherit deploy credentials just because both tasks happen during the same outage.

Evidence handling. Prefer outputs that preserve timestamps, links to source logs, and the distinction between an observed symptom and a proposed root cause. A neat narrative without traceable evidence makes the post-incident review harder.

Compatibility and setup. Check the listed agent support and install instructions against your actual client. A directory command is a starting point, not proof that a skill works with your repository or its current runtime.

Failure behavior. During an outage, a failed tool call should be visible. Repeated requests, unsolicited channel messages, and unapproved rollbacks can compound the incident.

Top agent skills for incident response

1. Incident Responder

Incident Responder is the first-hour option. Its directory entry describes structured triage, root cause analysis, mitigation steps, and a post-mortem draft following SRE runbook patterns. Give it the alert, the affected service, and a bounded set of logs. Ask for a timeline that labels observations and hypotheses separately. The on-call engineer still decides whether to run a mitigation.

Best for: Turning scattered incident evidence into a triage plan and a reviewable timeline.

Compatible with: Claude Code, Codex
Category: Engineering
Install: gh skill install ComposioHQ/awesome-claude-skills/incident-responder

2. Retry Handler

Retry Handler is worth opening when clients are making a bad situation worse. The directory describes exponential backoff, jitter, per-error policies, circuit breaker integration, and telemetry. Suppose a dependency is intermittently failing: an engineer can use the skill to inspect the client policy and draft a change. First identify errors that should never be retried, and cap the attempts. Otherwise, the proposed fix might just send more traffic into an already degraded service.

Best for: Reviewing retry policy when transient failures are amplifying an incident.

Compatible with: Claude Code, Codex, Cursor
Category: Orchestration
Install: gh skill install VoltAgent/awesome-agent-skills/retry-handler

3. Deployment Validator

Deployment Validator belongs at the change boundary, not at the start of triage. The entry covers environment variable completeness, migration status, health endpoint checks, rollback-plan review, and a stakeholder notification draft. Before a release, run the checklist against the proposed change. After a rollback, record what you checked and what changed. A plausible diagnosis alone is not authorization to deploy either one.

Best for: Checking a proposed fix or rollback plan before an authorized operator changes production.

Compatible with: Claude Code, Codex
Category: Engineering
Install: gh skill install VoltAgent/awesome-agent-skills/deployment-validator

4. Slack Integration

Slack Integration can recover the decision trail from a busy incident channel. It can search threads, retrieve messages, post to channels, and handle notifications, according to the directory. Use the read side first: find the last handoff and any decision to pause a rollout. Then draft an update that says who is affected, what is known, and when the next update is due. Keep channel posting behind human approval. Silence for a few minutes is preferable to a confident but stale status message.

Best for: Recovering the decision trail and preparing updates for the incident commander.

Compatible with: Claude Code, Cursor, Universal
Category: Automation
Install: gh skill install anthropics/knowledge-work-plugins/slack

How to choose

If you can install only one, start with Incident Responder for a consistent triage record. Add Slack Integration when handoffs across shifts are the problem, but separate read access from permission to post. Choose Retry Handler only when retry behavior is relevant to the failure mode. Use Deployment Validator when the response involves a release, a rollback, or a changed environment.

A practical sequence is: collect evidence, draft a hypothesis and timeline, have a human select a mitigation, validate the change, then send an approved update. Record what the agent actually saw and what an operator decided. The distinction matters more than whether the final timeline reads smoothly.

FAQ

Can an agent skill resolve a production incident on its own?
Not safely by default. A skill may prepare commands or recommend a rollback, but production changes should follow your access controls, incident roles, and approval process.

Should I install all four skills in one agent?
Not necessarily. A read-only triage agent and a separately authorized deployment workflow create a clearer boundary than one agent with every permission. Install only what the incident workflow needs.

Are the install commands verified against the latest upstream repositories?
The commands above are the values listed in the AgentNDX directory. Check the upstream repository and the installer supported by your agent before running them; directory metadata does not prove that a command is currently available.