Government Technology Review
Government ITLong read

Cybersecurity Incident Response Planning for Public Agencies

Agencies must align three separate regulatory frameworks to handle one incident correctly.

Features Editor · · 14 min read
Cover illustration for “Cybersecurity Incident Response Planning for Public Agencies”
Government IT · August 7, 2026 · 14 min read · 3,154 words

No single framework governs public agency incident response. In practice, agencies must satisfy multiple overlapping standards simultaneously, each with different origins, different scopes, and different compliance timelines. That layering is not bureaucratic redundancy. Each framework is solving a different piece of the problem, and the gaps between them are exactly where agencies get caught.

NIST SP 800-61 Revision 3, finalized in April 2025, is the foundational document and the first substantive update since 2012. The structural shift matters more than the revision number suggests. Instead of treating incident response as a standalone lifecycle, the revision maps IR recommendations directly to the six functions of the NIST Cybersecurity Framework 2.0: Govern, Identify, Protect, Detect, Respond, and Recover. Agencies are no longer expected to manage an IR plan as a separate IT security procedure. It is explicitly a component of enterprise-wide risk management governance now. Leadership owns it, not just the security team.

CISA's draft National Cyber Incident Response Plan, published December 16, 2024, updates the 2016 original and was developed with more than 150 experts from 66 organizations. It organizes federal coordination along four lines of effort: Asset Response, Threat Response, Intelligence Support, and Affected Entity Response. The draft defines an explicit participation path for non-federal stakeholders, positioning state and local agencies as active nodes in coordinated response rather than passive recipients of federal guidance.

CIRCIA, enacted in March 2022, mandates that covered entities report substantial cyber incidents within 72 hours and ransom payments within 24 hours. The final rule has been delayed to May 2026, but agencies that wait to design IR workflows around these timelines will find themselves engineering against a hard deadline with no runway. Retrofitting reporting requirements onto an existing IR process is considerably harder than building them in from the start.

CISA's Cybersecurity Performance Goals 2.0, released December 11, 2025, are voluntary but outcome-driven. They cover both IT and operational technology environments and function well as a practical gap analysis tool when run alongside mandatory requirements: less a compliance checklist, more an honest mirror.

NIST governs how you manage the incident internally. The NCIRP governs how you coordinate externally. CIRCIA governs what you are legally required to report and when. Three different dimensions, three different documents, one incident.

Diagram: Three Frameworks, Three Dimensions of One Incident. Visualizes: Visualize how three distinct regulatory frameworks each govern a different dimension of the same incident response obligation.

Phase 1 — Preparation: what agencies must have in place before any incident occurs

Preparation is where most agencies underinvest, and it is also the phase that determines how every subsequent phase performs. You are executing against whatever foundation you built before the alert fired.

An agency that arrives at a ransomware incident without a tested response plan, defined roles, or pre-negotiated vendor contracts will spend the first 24 to 48 hours making decisions that should have been made months earlier, under conditions no one would choose for making them. I have seen this happen. The forensic firm you call for the first time at 2 a.m. on a Saturday has leverage you do not want them to have.

The core deliverables are not complicated, but they require sustained institutional commitment to actually produce. An Incident Response Plan needs to be a written, approved, regularly tested document, not something that lives in a shared drive and gets dusted off at annual audits. It defines who declares an incident, who leads technical response, who handles communications, and who notifies legal and regulatory bodies. Pre-approved communication templates, covering internal notifications, press statements, and victim notification letters, belong in that same package. Agencies that draft these under pressure make errors, and errors during an active incident carry real costs downstream.

Vendor and mutual-aid agreements deserve specific attention, and this is where government procurement creates a genuine structural problem. Pre-negotiated contracts with IR retainer firms, forensic vendors, and peer agencies need to exist before something happens. The procurement system was not designed for emergency timelines. Work around it in advance. This means getting legal and procurement aligned on a retainer structure before anyone ever needs it, which is a harder organizational lift than it sounds.

Asset inventory and data classification are foundational in a way that is easy to underestimate until you are three hours into an incident and nobody can agree on whether a particular server holds PII. You cannot prioritize containment of systems you have not catalogued.

The staffing reality for government is stark. The GAO identified lack of staff as one of three primary challenges hindering federal agency IR preparedness, and state and local agencies are structurally understaffed even relative to federal ones. A large majority of state, local, tribal, and territorial organizations have fewer than five dedicated cybersecurity employees, according to KnowBe4. Plans that build in mutual aid agreements, third-party surge capacity, and explicit escalation paths to CISA and MS-ISAC have a materially better chance of holding.

Event logging gets treated as a detection-phase concern when it is actually a preparation-phase requirement. Twenty of the 23 CFO Act agencies failed to reach maturity level 3 of OMB's M-21-31 on event logging. Without adequate logging infrastructure in place before an incident, detection is impaired and forensics are crippled.

Tabletop exercises should be conducted regularly, and one per year is a floor, not a target. More importantly, these exercises need to include leadership, legal counsel, communications staff, and IT together in the same room. An exercise that only engages the security operations center does not test the decisions that actually matter when something real is happening. Compromised credentials were the entry point in 49% of ransomware attacks on state and local governments in 2024, which means user-level security awareness is a preparation-phase control, full stop.

NIST SP 800-61 Rev. 3 aligns preparation activities to the Govern and Protect CSF functions. Mapping preparation deliverables explicitly to those functions makes compliance documentation considerably easier and audit conversations less painful, which is a practical benefit worth capturing.

Phase 2 — Detection and analysis: moving from alert noise to confirmed incident

Tools generate alerts. Structured processes and trained analysts determine what those alerts mean. Agencies that invest in detection tooling without investing in the human and procedural layer end up with alert fatigue, missed signals, and delayed response.

The NCIRP treats monitoring, analysis, and detection as a unified set of activities rather than a binary trigger. That framing reflects how incidents actually unfold. Most significant incidents are not announced by a single, unambiguous alert. They emerge from a pattern of signals across multiple systems over time, and recognizing that pattern requires both decent tooling and analysts who know what they are looking for.

Common detection sources for public agencies include SIEM and SOAR tooling, which requires tuned rules and behavioral baseline data to function accurately; CISA's Einstein and Continuous Diagnostics and Mitigation programs for federal agencies; MS-ISAC alerts and threat intelligence feeds for state and local governments; and end-user reports. A trained employee who reports a suspicious email before clicking it is a detection mechanism, and in phishing and credential-theft scenarios, they are frequently the first one to fire.

The GAO's identified gap around event logging technical challenges hits hardest in this phase. Agencies that cannot correlate logs across systems cannot reconstruct attack timelines or confirm scope, and this is where underfunded logging infrastructure stops being a compliance shortfall and becomes an operational crisis.

Before moving to containment, the IR team needs answers to a specific set of questions: Which systems are affected, and what data do they hold? Is the attacker still active in the environment? What is the likely access vector? Does this event meet the threshold for mandatory reporting under CIRCIA or agency-specific requirements? These should be on a checklist that was built before the incident, not assembled during it.

Incident classification and severity tiering need to exist as a documented scale before an incident occurs. Without a defined severity framework, every incident either gets over-escalated or under-resourced.

Documentation starts here. Timestamps, evidence chain of custody, analyst notes. The temptation to move fast and document later is understandable; I have felt it. Resist it. The documentation captured during detection is what makes legal proceedings, insurance claims, and post-incident review function properly, and you cannot reconstruct it accurately after the fact.

The GAO's third identified gap, limitations in cyber threat information sharing, is most acute during this phase. Agencies that share indicators of compromise with CISA and MS-ISAC early benefit from pattern-matching across peer organizations. A signal that looks ambiguous in isolation can be conclusively identified when matched against incidents affecting other agencies in the same window.

Phase 3 — Containment: stopping the spread while preserving evidence

Containment is where the most consequential and often irreversible decisions get made under real-time pressure, which is precisely why pre-approved playbooks matter more here than anywhere else in the response lifecycle. An IR team deliberating about containment procedures while an attacker is actively moving through the environment will consistently arrive at those decisions too late.

Containment operates at two distinct levels. Short-term containment focuses on immediate isolation: network segmentation, account lockouts, blocking attacker-controlled IP ranges. The objective is stopping lateral movement before it extends further. Long-term containment stabilizes the environment so that operations can partially continue while eradication is prepared, typically involving system rebuilds in parallel or restoration from clean backups to isolated environments.

The backup targeting problem demands direct acknowledgment. In 2024, 99% of state and local government organizations hit by ransomware reported that attackers attempted to compromise their backups. Containment must therefore include immediate verification and isolation of backup infrastructure, not just production systems. An agency that contains its production environment while an attacker retains access to backup infrastructure has not contained the incident.

Evidence preservation is an area where government agencies face obligations that exceed those of most private entities. Public agencies have legal discovery requirements that can be triggered by litigation, oversight investigations, or regulatory proceedings. Forensic images of affected systems should be taken before any remediation action, full stop. Premature wiping, even when well-intentioned, has created legal complications in prior government incidents.

The Columbus, Ohio attack in July 2024 cost $7.3 million total and required 71 days to restore systems, roughly $102,817 per day of downtime. Decision-makers who underestimate how fast those costs accumulate tend to underinvest in early, aggressive containment.

Communication decisions during containment involve genuine tradeoffs that cannot be fully resolved in advance but can be anticipated. CIRCIA's 72-hour reporting clock starts at the point of reasonable belief, not confirmed attribution. Agencies need notification workflows ready before analysis is complete. At the same time, premature public disclosure can alert attackers who are still present in the environment. A pre-determined hold-and-notify protocol that balances these competing pressures, developed before any incident occurs with legal counsel involved, resolves most of these situations before they become conflicts under pressure.

The NCIRP's Asset Response and Threat Response lines of effort typically activate in parallel during containment. Law enforcement liaison contacts, whether FBI for federal entities or MS-ISAC and state fusion centers for state and local agencies, should be established before an incident. Introducing yourself to your FBI cyber contact for the first time while an attacker is in your network is a failure of preparation.

Phase 4 — Eradication: removing the threat completely before recovery begins

The single most common failure mode in incident response is declaring victory after containment and moving to recovery before the threat actor has been fully removed. The result is reinfection, extended downtime, compounding costs, and for public agencies, a second round of public accountability that is considerably harder to weather than the first.

Eradication requires a specific set of tasks completed in full before any restoration begins. The initial access vector needs to be identified and closed: patch the exploited vulnerability, reset compromised credentials, revoke stolen certificates. All malware, backdoors, and persistence mechanisms must be removed, which presupposes that thorough forensic analysis was conducted during and after containment. Privileged accounts need to be audited for unauthorized changes or additions. Software integrity and configurations should be verified against known-good baselines. Skipping any of these steps in the interest of speed is how agencies end up doing this twice, on a worse timeline the second time.

Nation-state and sophisticated criminal actors make eradication genuinely difficult in ways that routine ransomware does not. The PRC-associated actors who compromised more than 400 organizations through Microsoft SharePoint in 2025, including DHS and the Department of Energy, employed living-off-the-land techniques that blend into normal administrative activity. Standard malware scanning will not find them. Eradication in these scenarios requires behavioral analysis by experienced forensic professionals, and agencies that do not have that capability in-house need a retainer relationship with a firm that does. That relationship needs to exist before the incident.

Backup integrity remains a critical variable through eradication. Given that attackers targeting government agencies specifically attempt to compromise backup systems, eradication must confirm that backup infrastructure and recovery media are clean before any restoration begins. An agency that restores from a compromised backup has not recovered. It has restarted the incident under a different name.

Coordination with law enforcement during eradication requires a protocol established in advance, because the competing interests here are real and the tension does not resolve itself under pressure. Law enforcement may request that agencies delay certain eradication steps to preserve evidence or support attribution efforts. Agencies also have an obligation to restore services to the public. A pre-established protocol developed in consultation with legal counsel and law enforcement contacts resolves most of these situations before they become adversarial.

Every action taken during eradication should be logged: every system modified, every artifact removed, every decision made and by whom. This documentation becomes the basis for the forensic timeline used in lessons-learned analysis, litigation support, and regulatory reporting.

Phase 5 — Recovery: restoring operations without reintroducing the threat

Recovery is not a single event. It is a sequenced set of decisions about which systems to bring back first, in what order, and with what validation gates before each system returns to production. Agencies that treat recovery as a simultaneous restoration effort rather than a deliberate sequence tend to reintroduce the threat, extend their recovery timelines, and incur costs that were entirely avoidable.

The prioritization logic for government agencies follows a hierarchy that should be decided before any incident occurs, not negotiated during one. Life-safety and emergency services come first: 911 dispatch, hospitals, utilities for agencies with operational technology environments. Core government functions come second: payroll, benefits disbursement, court systems. Administrative and back-office systems come last. This sequence reflects both public safety obligations and the accountability expectations government agencies operate under. It is not primarily a technical decision. It belongs with leadership, informed by technical staff, and it needs to be documented before anyone is under pressure to make it.

The 2023 Dallas attack, which disrupted police, fire, and court systems simultaneously, illustrates what poor recovery sequencing costs beyond dollars. Some failures in incident response are embarrassing. Others carry public safety consequences that outlast the incident itself by years.

Before any system is reconnected to the network, it should be tested in an isolated environment against known-good configuration baselines. This is the control that prevents a multi-week recovery from becoming a multi-month one, and it is also the step that gets skipped most often when organizations are fatigued and under pressure to restore services. The fatigue is real. Skip the step anyway.

The cost trajectory of inadequate recovery planning is documented and has gotten worse. Mean recovery costs for state and local government ransomware incidents more than doubled, from $1.21 million in 2023 to $2.83 million in 2024, driven in significant part by attackers destroying backup infrastructure. Agencies without clean, tested, offline backups face the most expensive recovery paths available.

Communication obligations during recovery are substantial and easy to underestimate. Elected officials, oversight bodies, and the public expect ongoing updates. The 2025 PowerSchool breach exposed personal information of 62 million students and 9.6 million teachers; the Dallas attack exposed data belonging to 30,000 residents. Notification obligations at that scale require pre-built workflows. Improvising outreach to tens of thousands of affected individuals during active recovery is a failure of preparation that manifests as a communications crisis, and those two things require different resources to address.

Monitoring intensity should increase, not decrease, immediately after recovery. The period following restoration is a high-risk window for re-attack. Threat actors who retained access or simply observed the incident will attempt re-entry while the agency is still in a degraded operational posture. Enhanced logging, active threat hunting, and heightened analyst attention should continue for weeks, not days.

Diagram: Recovery Costs for State and Local Governments: 2023 vs. 2024. Visualizes: Show the stark year-over-year increase in mean ransomware recovery costs for state and local governments: $1.21 million in 2023 rising to $2.83 million in 2024 —…

Phase 6 — Post-incident review: turning the incident into institutional knowledge

The post-incident review is the phase most frequently skipped under resource pressure, and it is also the phase whose absence compounds every future incident's cost. An agency that moves from recovery directly back to normal operations without structured review has paid the full price of the incident and extracted none of its informational value.

Timing matters more than most agencies realize. The review should occur within two weeks of incident closure, while details are still fresh and the people who were involved are still accessible and willing to be honest about what went wrong. Reviews conducted months later, when staff have moved on or the memory of specific decisions has softened into institutional mythology, produce thinner analysis and weaker remediation commitments.

The review should produce specific, actionable outputs, not a narrative summary that gets filed and forgotten. A factual timeline of events from initial compromise through recovery belongs at the center of the document, built from the logs and documentation accumulated throughout the response. What the team got right should be documented explicitly, not assumed or left implicit. What failed or was absent, whether a missing playbook, an untested backup, an unclear authority structure, needs to be named with specificity, because vague findings produce vague remediation. Each identified gap should be assigned to an owner with a target remediation date. Without ownership and a deadline, findings persist as acknowledged problems rather than resolved ones, sometimes for years.

The review should also assess whether the agency met its regulatory reporting obligations and on what timeline. If CIRCIA's 72-hour or 24-hour windows were not met, the post-incident review is where to identify why and redesign the workflow before the next incident creates the same problem under worse circumstances.

Institutional memory is a genuine vulnerability for government agencies, and one that does not get enough attention in security planning. Staff turnover in cybersecurity roles, where demand consistently exceeds supply, means that the organization's knowledge of past incidents walks out the door with departing employees unless it has been systematically documented. The post-incident review is how the agency retains what it paid, sometimes very dearly, to learn.

The review's outputs should feed directly back into Phase 1: updated playbooks, revised tabletop exercise scenarios, corrected asset inventories, amended communication templates. The feedback loop from Phase 6 to Phase 1 is what separates agencies that get better between incidents from agencies that simply repeat them, at escalating cost.

Sources

  1. cisa.gov
  2. cisa.gov
  3. cisa.gov
  4. csrc.nist.gov
  5. itif.org
Filed underGovernment IT

More in Government IT