Tuesday, September 22, 2026

Rehearse a Data Center Crisis Against 10,000 Simulated Devices

 
Data centers are getting more public scrutiny than they have in years. Whatever
your view of that debate, one thing is clear: a data center that does get built should
work on day one. It should not fail its first crisis, and it should not be over-built to 
cover for what nobody tested.

Most of what goes wrong in a new facility is not the building. It is the network and 
the tools watching it: the detection rule that never fires, the monitoring system that 
drowns in its first alarm storm, the configuration push that stalls at device four 
thousand. None of those can be tested on the production network, and none of 
them show up at the scale of a lab.

Rehearse it, don't hope for it

 
MIMIC Simulator assists in hardening a data center by giving you a data center's 
worth of network devices that cost nothing to break -- thousands of simulated 
switches, routers and servers, each on its own IP address, answering SNMP, 
exporting NetFlow, IPFIX and sFlow, and accepting CLI logins over Telnet and 
SSH, all running on a server rather than in racks.

Point your security, monitoring and configuration tools at it, and you can run the 
attack, the outage and the bad change before they happen for real. The resilience 
comes from the rehearsal.

If you read our August post, you have seen how to lay out a ten-thousand-device 
topology in minutes. This is what to do with it.

1. Rehearse the attack

A threat-detection platform is only as good as the rules you have tuned in it,
and the honest way to tune a rule is to show it the thing it is meant to catch.
On a production network you cannot schedule a DDoS or a port scan. On a
simulated one, you can, many times.

The MIMIC NetFlow Simulator exports flow records from every simulated device, 
and you control every field in them: source and destination addresses, protocols, 
ports, packet and byte counts, export intervals. That lets you produce, on demand 
and as often as you like:

  - a DDoS over TCP or UDP against one of your own addresses;
  - a port scan or reconnaissance sweep across a subnet;
  - brute-force login attempts;
  - large outbound flows to an unknown destination at an odd hour -- the
    shape of data exfiltration;
  - and the quiet versions of each, which are the ones worth testing: the
    low-volume exfiltration, the scan hidden in normal traffic.

Then watch what your security application does with it. Does the alert fire?
Does it fire on the stealthy version, or only the loud one? Does the ordinary
application traffic running alongside it (simulated by Cisco AVC application
flows) trigger false positives? Adjust the threshold, replay the same traffic,
and compare. Because the traffic is simulated, the second run is identical to
the first, which is what makes the comparison mean something.

 
Video: "MIMIC NetFlow Simulator and ElastiFlow NetObserv"

The video above shows this against ElastiFlow NetObserv: MIMIC generates 
DDoS, port-scan, reconnaissance and brute-force traffic, and ElastiFlow picks it 
up. The full walkthrough is in our ElastiFlow post. Seceon used the same 
approach to develop its data-center security platform, with simulated switches on 
every rack exporting NetFlow into its detection engine.

2. Rehearse the outage

The first real crisis in a new data center is also the first time the monitoring 
system sees one. That is a poor time to learn that its dashboards slow to a crawl 
at ten thousand devices, or that an alarm storm buries the one alarm that 
mattered.

With the MIMIC SNMP Simulator, every simulated device carries a live MIB you 
can change while it runs. So you can stage the outage itself:

  - take interfaces down and bring them back, one or a thousand at a time;
  - send the trap storm that a failing core switch produces;
  - raise error counters and utilization on a link until it saturates;
  - stop devices outright, and see how long it takes your system to notice.

Two things come out of that. The first is a measured answer to "does our 
monitoring cope?", taken at the scale you will actually run. The second is a 
trained team. Operators can work through a staged crisis on a network that looks 
exactly like the real one, without anything real being at risk. Pepco Holdings did 
exactly this -- see Identifying and preparing for Crisis Situations --  running 
10,000 simulated devices to prepare its operations center for worst-case events.

3. Rehearse the change

Configuration management tools are tested on a handful of devices and then
trusted with thousands. The MIMIC IOS Simulator gives your chosen tool
thousands of devices to log in to over Telnet or SSH, each answering the CLI.
Run your change against the whole fleet and find out how long it takes, what
happens when devices stop answering halfway through the run, and whether your
tool reports those failures or quietly skips them. Vistara used MIMIC's SNMP 
and Telnet/SSH simulation to test its platform at enterprise scale.

Build it right

Whether a data center should be built is not a question a simulator can answer. 
How well it works once it is built is, and the answer depends on how much of the 
first crisis you have already seen.

Every rehearsal above runs on your schedule, repeats identically, and costs 
nothing when it goes wrong -- which is the point. If there is a scenario your
team has been meaning to test and never could, tell us what it looks like.

Customers have used MIMIC for:
  - data-center security:  Seceon
  - crisis preparation:  Pepco Holdings
  - configuration management:  Vistara
  - data-center infrastructure management:  Device42, Graphical Networks