0xmichalis

How We Built Our Master Security Incident Response Playbook (and Why You Need One Too)

I recently had a conversation with a friend and former colleague, and we reminisced about how we had organized the company’s ecosystem back when we worked together (he still works there) so that we could activate our playbooks quickly and clearly. The most important thing we did together was to structure and implement the Master Security Incident Response Playbook, which, even now, is used as the main reference for their incidents. Since the way we built it and its usefulness might also benefit others, I decided to share a more general version of it on my blog. I truly hope you find it useful. :)

masterplaybook

This Master Security Incident Response Playbook serves as the central guide for the Incident Response (IR) team. Its primary role is to direct the initial steps: from detection and assessment to the precise categorization of each incident.

Once an incident is categorized (e.g., as ransomware, data leak, or phishing), the master playbook triggers and refers to the corresponding specialized playbook. These specialized playbooks contain the detailed, technical steps required for the effective handling of the specific threat.

The original version of this playbook was based on real-world events from service providers. Still, it has been fully adapted, with sensitive information removed and the content modernized to meet today’s security requirements.

The Markdown format was chosen for practical reasons:

As an example, the playbook details the case of a data leak, but its structure is designed as a master copy from which all other specialized playbooks can be adapted.

Finally, for completeness and effectiveness, the playbook includes references to relevant tables, policies, procedures, and guidelines. This ensures that all involved parties have access to the necessary information both during preparation and testing, as well as when activating the playbook in a real incident. If you’d like to download the full playbook, feel free to grab it from my GitHub repo: Github Link



Master Security Incident Response Playbook

Guide for Comprehensive Cybersecurity Incident Management


Table of Contents


1. Executive Summary

This playbook offers a practical, structured approach for the rapid and coordinated response to cybersecurity incidents. It is based on industry best practices and can be easily adapted to the needs of any organization, serving as a guide for all specialized playbooks.

Objectives:


2. Incident Classification Matrix

2.1 Severity Levels

The severity level is defined by the Organization based on the regulatory framework and management requirements. Critical systems are those internally designated based on Business Impact Analysis, DORA criticality, and GDPR criticality. (Reference to Critical Systems Registry).

CRITICAL (P1): Critical systems have been breached, confirmed data leak, regulatory notification required

HIGH (P2): Significant breach, possible data leak, service disruption affecting customers.

MEDIUM (P3): Limited breach, no confirmed data leak

LOW (P4): Suspicious activity, possible incident under investigation

2.2 Incident Categories

For each detected incident, the IR team must immediately categorize it into one of the following categories to follow the appropriate response procedures:

Malware / Ransomware
Indicators: Ransom messages, mass file encryption, and malware detection by security tools.
Immediate actions: Isolate affected systems, preserve evidence, notify IR team, initiate containment, start relevant playbook (Reference to Malware/Ransomware Playbook)

Data Leak/Exfiltration
Indicators: Unexplained large data transfers, sending sensitive files outside the Organization, DLP system alerts.
Immediate actions: Identify source of leak, isolate related accounts/systems, preserve logs, notify DPO/legal, start relevant playbook (Reference to Data Leak/Exfiltration Playbook)

Unauthorized Access
Indicators: Access to systems with stolen credentials, unusual account activity, and SIEM alerts.
Immediate actions: Disable/reset credentials, check logs for lateral movement, isolate affected accounts, start relevant playbook (Reference to Unauthorized Access Playbook)

Denial of Service (DoS/DDoS)
Indicators: Service unavailability, excessive requests to servers, and monitoring alerts.
Immediate actions: Activate DDoS mitigation, notify provider, log attacks, communicate with affected stakeholders, start relevant playbook (Reference to DoS/DDoS Playbook)

Phishing / Social Engineering
Indicators: Reports of suspicious emails, users providing credentials, and detection of malicious links.
Immediate actions: Notify affected users, reset credentials, analyze email headers, notify awareness team, start relevant playbook (Reference to Phishing / Social Engineering Playbook)

Insider Threat
Indicators: Unusual access or data exfiltration by internal users, reports of suspicious behavior.
Immediate actions: Access audit, isolate account, preserve evidence, notify HR/legal, start relevant playbook (Reference to Insider Threat Playbook)

Third Party / Supply Chain
Indicators: Notification of compromise from partner, detection of malicious activity via external services, vulnerability feed alerts.
Immediate actions: Communicate with third party, isolate affected integrations, assess impact, notify management, start relevant playbook (Reference to Third Party / Supply Chain Playbook)

Physical Security Breach
Indicators: Theft or loss of equipment, unauthorized entry to premises, reports of vandalism.
Immediate actions: Notify physical security, preserve CCTV/logs, isolate affected systems, activate disaster recovery if required, start relevant playbook (Reference to Physical Security Breach Playbook)


3. Team Roles & Responsibilities

Effective incident management relies on a clear division of roles and immediate collaboration among Incident Response (IR) team members. Each role has specific responsibilities and actions, while the Incident Commander (IC) ensures overall coordination and decision-making.

The IR team consists of:

Coordination among roles is achieved through regular communication, clear assignment of responsibilities, and documentation of all actions. The Incident Commander is the sole point of decision-making and communication with management, while each lead ensures timely updates and collaboration with other team members. (Reference to Crisis Management and Internal Communication Form).

3.1 Incident Commander (IC)

The Incident Commander (IC) is responsible for the overall coordination of incident management. They lead the IR team from the moment the playbook is activated until full recovery and incident closure.

Key Responsibilities:

Immediate actions upon activation:

> Note: The IC remains the sole point of decision-making and communication with management throughout the incident, ensuring a unified strategy and avoiding confusion.

3.2 Security Lead

The Security Lead is responsible for the technical management of the incident and threat assessment. They act as the primary technical coordinator of the IR team, ensuring proper analysis, documentation, and preservation of evidence.

Key Responsibilities:

Immediate actions upon activation:

> Note: The Security Lead ensures that the technical management of the incident is conducted methodically, with integrity, and by organizational policies.

3.3 IT Operations Lead

The IT Operations Lead is responsible for managing the technical infrastructure during the incident, ensuring the isolation, recovery, and continuity of critical services.

Key Responsibilities:

Immediate actions upon activation:

> Note: The IT Operations Lead ensures that technical actions are performed securely and minimize business impact.

3.4 Communications Lead

The Communications Lead is responsible for managing communication, both internally and externally, during the incident.

Key Responsibilities:

Immediate actions upon activation:

> Note: The Communications Lead ensures that communication is timely, accurate, and consistent with organizational policies.

3.5 Legal/Compliance Officer

The Legal/Compliance Officer is responsible for managing the legal and regulatory aspects of the incident.

Key Responsibilities:

Immediate actions upon activation:

> Note: The Legal/Compliance Officer ensures that all actions comply with the applicable legal and regulatory framework.


4. Communication Templates

4.1 Initial Notification Template

SUBJECT: [SEVERITY] Security Incident - [INCIDENT-ID] - [BRIEF DESCRIPTION]

INCIDENT DETAILS:
- Incident ID: [AUTO-GENERATED] - If not available, simply use the date in DDMMYYYY format
- Detection time: [TIMESTAMP]
- Severity level: [P1/P2/P3/P4]
- Affected systems: [LIST] - Does not need to be exhaustive, only highlight the most critical ones. Update with new data in each notification.
- Initial assessment: [BRIEF DESCRIPTION]
- Incident Commander: [NAME/CONTACT]

IMMEDIATE ACTIONS [SUMMARY]:
- [ACTION 1]
- [ACTION 2]
- [ACTION 3]

NEXT STEPS [SUMMARY]:
- [PLANNED ACTION - TIME]
- [PLANNED ACTION - TIME]

Next update: [In minutes]

4.2 Executive Briefing Template

EXECUTIVE INCIDENT BRIEFING - [INCIDENT-ID]

SITUATION OVERVIEW:
- What happened: [SIMPLE, NON-TECHNICAL DESCRIPTION]
- When detected: [TIMESTAMP]
- Current status: [CONTAINED/INVESTIGATING/RECOVERING]
- Business impact: [QUANTIFIED] - If quantification is not possible, refer to criticality level P1/P2

RESPONSE ACTIONS:
- Immediate containment: [ACTIONS]
- Investigation progress: [KEY FINDINGS]
- Recovery timeline: [ESTIMATE]

LEGAL/REGULATORY:
- Notification required: [YES/NO - TIME]
- Impact on customers: [DESCRIPTION]
- Media/PR: [ASSESSMENT]

RESOURCE NEEDS:
- Additional personnel: [IF NEEDED]
- External support: [VENDORS/CONSULTANTS]

Next update: [In minutes]

5. Technical Response Procedures

This section describes the basic technical steps the IR team must follow from the moment an incident is detected until the threat is wholly eradicated. The process is structured to ensure speed, accuracy, and preservation of evidence.

5.1 Initial Response Checklist

During the initial phase, the IR team follows a specific checklist to ensure all critical steps are covered:

5.2 Containment Strategies

Containment aims to limit the spread of the incident and protect critical systems. It is divided into short-term and long-term actions:

Short-term containment:

Long-term containment:

5.3 Eradication Procedures

Once containment is achieved, the IR team proceeds with the complete eradication of the threat, depending on the type of incident:

Malware Removal:

  1. Identify all affected systems (using IOCs).
  2. Apply anti-malware tools to remove malicious software.
  3. Manually remove persistent threats (e.g., rootkits, backdoors).
  4. Clean registry and file system remnants.
  5. Confirm complete removal through repeated checks.

Account Compromise:

  1. Mandatory password change for affected accounts.
  2. Check and revoke suspicious tokens/sessions.
  3. Audit permissions and groups to identify unauthorized changes.
  4. Enable additional authentication (e.g., MFA).
  5. Enhanced monitoring for new unauthorized access.

System Compromise:

  1. Restore systems from backups.
  2. Apply all necessary patches and updates.
  3. Harden security settings (system hardening).
  4. Restore data only from verified backups.
  5. Security testing and validation before restoring to production.

> Note: All actions must be thoroughly documented, and chain of custody procedures for evidence must be followed.


6. Evidence Collection & Preservation

The collection and preservation of evidence is a critical process for forensic analysis, impact assessment, and incident documentation. Proper evidence management ensures its legal validity and integrity.

6.1 Digital Evidence Management

Chain of Custody:
The chain of custody process ensures the integrity and reliability of evidence:

Types of Evidence:
The IR team collects different types of evidence depending on the nature of the incident:

6.2 Evidence Collection Commands

The IR team uses specific commands to collect evidence from different environments. Some basic examples are listed below; each playbook contains a particular set of commands and actions that need to be executed.

Windows Systems:

# Memory dump
DumpIt.exe /output C:\evidence\memory.dmp

# System information
systeminfo > C:\evidence\systeminfo.txt

# Network connections
netstat -ano > C:\evidence
etstat.txt

# Running processes
tasklist /v > C:\evidence\processes.txt

# Event logs export
wevtutil epl System C:\evidence\system.evtx
wevtutil epl Security C:\evidence\security.evtx
wevtutil epl Application C:\evidencepplication.evtx

Linux Systems:

# Memory dump
dd if=/dev/mem of=/tmp/evidence/memory.dump

# System information
uname -a > /tmp/evidence/system_info.txt
cat /proc/version >> /tmp/evidence/system_info.txt

# Network connections
netstat -tulpn > /tmp/evidence/network_connections.txt

# Running processes
ps aux > /tmp/evidence/processes.txt

# Log files collection
cp -r /var/log/* /tmp/evidence/logs/

Network & Infrastructure:

# Packet capture
tcpdump -i eth0 -w /tmp/evidence/network_capture.pcap

# Continuous capture with rotation
tcpdump -i eth0 -w /tmp/evidence/capture_%Y%m%d_%H%M%S.pcap -G 3600

# DNS query logs
dig @dns_server example.com > /tmp/evidence/dns_queries.txt

Additional Commands:

Firewall logs Can Be Exported from the management console or collected from log files.
Proxy logs: Export from the proxy server or web security gateway.
SIEM data: Export relevant events and alerts from the SIEM platform.

> Note: All commands must be executed with appropriate permissions and documented with timestamps. Evidence must be stored in a secure location with hash values calculated to confirm integrity.


7. Recovery & Business Continuity (Reference to BCP)

The recovery and business continuity phase ensures the safe and controlled return of services to regular operation, minimizing business impact and preventing recurring incidents.

7.1 Recovery Planning

Recovery planning prioritizes systems and services, ensuring that the most critical ones for the Organization's operation are restored first.

Recovery Priorities:

  1. Customer-facing services: Applications and services used by customers.
  2. Internal operations: Support systems for staff and daily operations.
  3. Administrative systems: Management, reporting, HR systems, etc.

Recovery Verification:
After each system is restored, specific verification steps are followed:

7.2 Business Continuity Activation

In cases where recovery is delayed or the availability of critical services cannot be ensured, business continuity plans (BCPs) are activated.

Activation Triggers:

Business Continuity Procedures:

  1. Activate alternative sites: Transfer operations to disaster recovery or backup sites.
  2. Manual workarounds: Implement manual procedures for critical operations.
  3. Redirect customer traffic: Redirect customers to available services or alternative solutions.
  4. Inform stakeholders: Continuous updates to management, customers, and partners on the situation and next steps.
  5. Continuous service monitoring: Intensive monitoring of service status and alternative solutions.

> Note: All recovery and business continuity actions must be thoroughly documented and communicated to relevant stakeholders.


8. Post-Incident Actions

The review and improvement phase ensures that every incident is utilized as a learning opportunity and strengthens the Organization's resilience. Systematic analysis and implementation of corrective actions reduce the likelihood of recurrence and improve overall response.

8.1 Lessons Learned

After the incident is closed, a lessons learned meeting is held with the participation of all involved teams. The meeting agenda includes:

  1. Timeline review: Presentation and analysis of the chronological sequence of events.
  2. Response effectiveness evaluation: Discussion of what worked well and what did not in incident management.
  3. Communication evaluation: Review of internal and external communication.
  4. Identification of technical gaps: Recording of technical weaknesses or deficiencies identified.
  5. Improvement proposals: Collection of ideas and suggestions for future enhancement.
  6. Assessment of training needs: Evaluation of whether additional training or awareness is required.
  7. Assignment of action items: Definition of responsible parties and timelines for corrective actions.

Required documentation:

  1. Complete and documented incident timeline.
  2. Root cause analysis with analysis of causes.
  3. Impact assessment (financial, operational, reputational).
  4. Detailed record of actions taken.
  5. List and description of evidence collected.
  6. Record of regulatory notifications made.
  7. Consolidated improvement proposals and lessons learned.

8.2 Implementation of Improvements

Improvements resulting from the lessons learned process are implemented at three levels:

Technical improvements:

  1. Strengthening existing controls and security mechanisms.
  2. Improving monitoring and detection (e.g., new rules, alerts).
  3. Changes in system architecture to reduce risk.
  4. Hardening of critical systems.
  5. Upgrading or replacing security tools.

Procedural improvements:

  1. Updating and revising the playbook and IR procedures.
  2. Strengthening and updating training and awareness programs.
  3. Improving communication between teams and stakeholders.
  4. Updating the escalation matrix and communication plans.
  5. Strengthening collaboration and relationships with vendors and third parties.

Organizational improvements:

  1. Changes in staffing and role allocation.
  2. Skill development and professional growth programs.
  3. Budget review to cover new needs.
  4. Updating security policies and procedures.
  5. Improving governance and overall security governance.

> Note: All improvements must be documented, their implementation monitored, and reviewed regularly.


9. Appendices

9.1 Contact Information

Internal Contacts (Primary and Backup):

External Contacts:

9.2 Regulatory Notification Requirements

GDPR (EU):

DORA - NIS 2 - EETT: