Break-fix

When two vendors each say it is the other one.

Laptops that fall to the quarantine VLAN after a certificate renewal, a VPN that drops large transfers after an ISP change, an application that breaks when a service account password rotates, a historian that stopped collecting at 02:00. We put a capture on the wire and let the packets settle it. Most integration faults are a timeout, an MTU, a byte order or a certificate, and all of them are visible if you look.

Why it matters

When two vendors each point at the other, the service, the building or the plant stays degraded while the argument runs. A capture ends the argument, because the fault is in the conversation and the conversation is on the wire.

The other risk is a fix that is really a guess, and breaks something else. We prove the cause before we change anything, and we write down what we changed and why.

What it covers

Network and security faults
Duplicate IPs, MTU and fragmentation black holes, spanning-tree events, asymmetric routing through stateful firewalls, VPN and NAT faults, 802.1X and RADIUS failures, multicast and IGMP snooping, and time-sync drift.
Identity and certificate faults
Kerberos and LDAP bind failures, a certificate chain that changed under a renewal, a service account whose password rotated, and config drift from a change nobody recorded.
Industrial protocol faults
Packet capture and dissection for Modbus, DNP3, IEC 60870-5-104, EtherNet/IP and OPC UA: polling rates and timeouts, DNP3 unsolicited responses and class 0/1/2/3 polls, Modbus register maps and byte/word order, OPC DA and DCOM versus OPC UA certificates and security policies.

How it runs

  1. Capture

    A mirror port on the conversation that is failing, so we are looking at what actually happens rather than what the manual says should.

  2. Find the cause

    Read the capture, the logs and the config diffs until the fault is named, not guessed. The fault is in the conversation, so we listen to the conversation.

  3. Fix in a window

    The change applied inside an agreed window, with a configuration backup taken before and after.

  4. Write it down and watch it

    A short note for the site file saying what changed and why, and where it helps, a monitoring rule so it is caught earlier next time. Fixed, written down, and watched for next time.

A page of what you are handed

ARTA CYBER Representative deliverable — client and site redacted.

Root-cause note — sample

Staff laptops failing 802.1X after the RADIUS certificate renewal

Environment: client environment (redacted) · Symptom: from the morning after a certificate renewal, some laptops land on the quarantine VLAN while the rest connect normally.

Cause. The RADIUS server certificate was renewed from a new intermediate CA. Laptops still on the older wireless profile trusted only the previous intermediate, so they rejected the new chain during EAP-TLS and never sent their own certificate. Capture radius-eap.pcapng shows those clients ending the TLS handshake with an unknown_ca alert straight after the server’s Certificate message; laptops on the newer profile complete it.

Fix. Pushed an updated wireless profile that trusts the issuing root and names the RADIUS server, rather than pinning one intermediate, through the existing device management with site IT’s agreement. Verified: the affected laptops authenticate and land on the staff VLAN, and the RADIUS log shows EAP-TLS success for each device that had been failing.

Caught earlier next time. Added a NOC rule to alert when EAP-TLS rejects rise above their normal level for more than 10 minutes, and a change-checklist step to test the profile against any renewed RADIUS certificate before it goes live.

What you are left holding

  • Root-cause statement with the evidence: captures, logs, config diffs.
  • The fix, applied inside an agreed window.
  • Before and after configuration backups.
  • A short note for the site file: what changed and why.
  • Where it fits, a monitoring rule so it is caught earlier next time.
OT Asset Simulator

When a fault is safe to reproduce, we rebuild it off-site in the OT Asset Simulator, virtual Modbus TCP or DNP3 devices standing in for yours, so we can try the fix somewhere that is not your live plant before we touch the real one.

OT Asset Simulator

Questions we are asked first

Something is degraded right now. How fast can you look?

The first move is a capture, which we can often set up remotely or talk your on-site staff through within the call. What we will not do is apply a guessed fix to a live service or process to look fast; we find the cause first.

What should we send you before you arrive?

A packet capture on the affected conversation if you can take one, the relevant config exports, and a note of anything that changed recently, including firmware, passwords and switch updates. There is a short checklist we can send you.

Will the fix break something else?

That is why we prove the cause with evidence before we change anything, take a backup before and after, and write down exactly what changed. A change we cannot explain is a change we do not make.

Tell us what is bothering you.

An email is enough to start with. A scoping call is free and there is nothing to commit to, and where we are not the right people we will say so and point you at someone who is.

Book a scoping call

A first call about one site. No charge, and nothing to commit to.

Ask about a Site Assessment

Our named next step: we come to one site and hand you an assessment you can act on.