
Debugging PFC Pauses and ECN Marks in a RoCEv2 AI Fabric
Why PFC Pauses Signal a Misconfigured RoCEv2 Fabric
Ethernet didn't win the AI cluster by getting faster. It won because 800G optics got cheap enough that working around Ethernet's congestion behavior stopped being the deciding cost. The thing you're working around is easy to state: RoCEv2 carries a transport built for a fabric with credit-based, per-link flow control, and Ethernet has no credits. It has pause frames and congestion marks. This post walks through how to debug both — the counters to read first, how to tell a PFC pause storm apart from an ECN marking failure, and the threshold ordering that keeps PFC as a backstop instead of a steady-state mechanism.
Which means the reliability of a training job now rests on two mechanisms most teams configure once, from a vendor cookbook, and never revisit until something stalls. Priority Flow Control (PFC) stops traffic. Explicit Congestion Notification (ECN) tells senders to slow down. Get the ordering wrong and the fabric pauses constantly while every sender keeps blasting at line rate, because nothing ever told the sender to ease off.
The position this post takes: if PFC is doing steady-state work in your fabric, your fabric is misconfigured. PFC should be a rare backstop that fires in the tail. ECN plus DCQCN is the mechanism that is supposed to keep the workload running. When rx_prio<n>_pause counters climb in lockstep with throughput, you have a threshold-ordering bug, not a bandwidth shortage.
How PFC Pause Frames and ECN Marking Interact in RoCEv2
PFC pause frames and per-priority buffer thresholds
PFC lives in IEEE 802.1Qbb. It extends the classic Ethernet PAUSE frame — MAC control opcode 0x0001 — with a new opcode and an 8-bit priority enable vector. One frame can pause up to eight priorities independently, and the pause duration field is 16 bits, counted in units of 512 bit times.
The trigger is a per-priority ingress buffer watermark. Cross the Xoff threshold for a given priority and the switch sends a PFC frame upstream asking the sender to stop that priority. Traffic resumes once the buffer drains below Xon. This is why PFC misconfiguration hurts so much: the pause is hop-local and priority-scoped. The upstream device has no idea which flow filled the buffer. It stops all of them.
It's also why PFC needs headroom. Between the buffer crossing Xoff and the pause landing on an upstream port that actually stops transmitting, bytes are still in flight. At 400 Gb/s, a 10 µs worst-case reaction window is roughly 500 KB of headroom per port — an estimate, but the right order of magnitude, and it explains why 800G line cards spend so much silicon on buffering.
ECN marking, CNP generation, and the DCQCN control loop
ECN is RFC 3168, and it occupies the low two bits of the IPv4 TOS byte (IPv6 traffic class). The codepoints are 00 Not-ECT, 10 ECT(0), 01 ECT(1), and 11 CE for Congestion Experienced.
RoCEv2 encapsulates InfiniBand transport over UDP destination port 4791, and the loop runs like this:
- A switch queue exceeds its ECN minimum threshold. Between min and max it marks probabilistically; above max it marks everything.
- The receiving host sees a CE-marked RoCE packet and generates a CNP (Congestion Notification Packet) — a RoCEv2 packet with its own transport opcode — back to the sender. Receivers coalesce CNPs with a minimum interval, so a burst of marked packets doesn't turn into a burst of CNPs.
- The sender runs DCQCN (Zhu et al., SIGCOMM 2015): on CNP it cuts its rate by a factor derived from a running
alpha, then recovers through additive increase, hyperactive increase while in fast recovery, and a fast-recovery state machine that also drivesalphaback toward zero.
Step 3 is the part that matters. It changes the sender's rate. PFC changes nothing about the sender's rate. That asymmetry is the whole story.
Why ECN marking and PFC thresholds fight each other
ECN and PFC are both triggered by queue depth, so the thresholds have to be ordered: ECN min < ECN max < PFC Xoff. Put ECN's thresholds above PFC's Xoff and the queue pauses before it ever gets marked. The sender is told to stop, never told to slow down, and resumes at the same rate the instant the pause expires. You get oscillation instead of a control loop.
There's a second, sneakier failure: ECN marking only happens on packets that carry an ECT codepoint. If a host egresses RoCE with the ECN bits set to 00, the switch sees Not-ECT and, on most implementations, won't mark it. Your only backpressure is PFC. That is a host-side configuration bug wearing the costume of a switch-side congestion bug, and it's worth ruling out before you touch a single switch threshold.
What a PFC Pause Storm Looks Like in Telemetry
Counters worth reading first on NICs and switches
Counter names vary by driver and vendor — check your own ethtool -S output rather than trusting a blog. The shape is what matters.
| What you want to know | Typical counter | Where |
|---|---|---|
| PFC pause frames received, per priority | rx_prio<n>_pause | NIC, ethtool -S |
| Cumulative pause duration, per priority | rx_prio<n>_pause_duration | NIC, ethtool -S |
| PFC pause frames sent, per priority | tx_prio<n>_pause | NIC, ethtool -S |
| CE-marked RoCE packets arriving | np_ecn_marked_roce_packets | NIC (mlx5) |
| CNPs generated | np_cnp_sent | NIC (mlx5) |
| PFC RX/TX per port per priority | show pfc counters | SONiC switch |
| Queue depth and watermarks | show queue counters, show queue watermark | SONiC switch |
| PFC state and watchdog events | show interface priority-flow-control | NX-OS switch |
Pause-duration units differ between drivers. Read the counter as a relative signal over time, not an absolute duration, unless the driver docs say otherwise.
Separating PFC pauses from ECN marks by timing and destination
Two structural facts make this easier than it sounds:
- PFC is hop-local. A pause counter increments on the link whose receiver buffer filled. It tells you where congestion materialized, not where it came from.
- CNPs are end-to-end. A CNP travels from the receiver back to the sender across the fabric, so
np_cnp_sentincrements on the receiver's NIC and the sender's rate drops accordingly.
So the diagnostic question becomes: on a link that's pausing, is the far end also emitting CNPs? If rx_prio3_pause climbs on a NIC while np_cnp_sent on that host stays flat, the packets arriving aren't being marked. Either the ECN threshold sits above the Xoff threshold, ECN isn't enabled on that queue, or the flows are Not-ECT. Three different fixes, all of them invisible in the pause counter.
A Practical Diagnostic Sequence for PFC and ECN Issues
Step 1 — baseline counters before the job runs
Snapshot counters and configuration before the workload, so every delta has a denominator.
# counter snapshot before the job
ethtool -S eth0 | grep -Ei 'pause|pfc|ecn|cnp|discard' | sort > /tmp/eth0.pre
## is this NIC's RoCE using an ECT codepoint at all?
cma_roce_tos -d mlx5_0
## NIC-side PFC and trust config
mlnx_qos -i eth0
## what does the switch advertise to us over DCBX/LLDP?
lldptool -t -i eth0 -V PFCExample output shape (illustrative, not a capture from a specific site):
$ cma_roce_tos -d mlx5_0
106
$ lldptool -t -i eth0 -V PFC
PFC-Version: 0
Enabled Priorities: 3
106 is 0x6A: DSCP 26 shifted left two bits, plus 10 for ECT(0). If cma_roce_tos returns 0, stop here. Nothing downstream can mark your traffic, and every congestion event will be resolved by PFC alone.
Step 2 — watch PFC and ECN counters in real time during the workload
ethtool -S is a wall of text. A small Node script turns it into a per-second delta stream, which is much easier to line up against a training log.
import { execFile } from "node:child_process";
const run = promisify(execFile);
const IFACE = process.argv[2] ?? "eth0";
const KEYS = /(pause|pfc|ecn|cnp|discard)/i;
function parse(text) {
const out = {};
for (const line of text.split("\n")) {
const m = line.match(/^\s*(\S+?):\s+(\d+)\s*$/);
if (m && KEYS.test(m[1])) out[m[1]] = Number(m[2]);
}
return out;
}
async function sample() {
const { stdout } = await run("ethtool", ["-S", IFACE], { maxBuffer: 1 << 22 });
return { t: Date.now(), c: parse(stdout) };
}
let prev = await sample();
setInterval(async () => {
const cur = await sample();
const dt = (cur.t - prev.t) / 1000;
const rows = Object.keys(cur.c)
.map((k) => [k, (cur.c[k] ?? 0) - (prev.c[k] ?? 0)])
.filter(([, d]) => d > 0)
.map(([k, d]) => [k, d, (d / dt).toFixed(1)]);
if (rows.length) {
console.log("\n== " + new Date(cur.t).toISOString());
for (const [k, d, r] of rows.sort((a, b) => b[1] - a[1])) {
console.log(k.padEnd(34) + "+" + d + " (" + r + "/s)");
}
}
prev = cur;
}, 1000);Run it with node watch-congestion.mjs eth0 (top-level await needs ESM, so keep the .mjs extension). Output shape:
== 2025-06-18T14:02:11.004Z
rx_prio3_pause +482 (482.0/s)
rx_prio3_pause_duration +15424 (15424.0/s)
np_ecn_marked_roce_packets +311 (311.0/s)
np_cnp_sent +298 (298.0/s)
Pauses and CNPs rising together is the healthy case. ECN is firing and PFC is catching the overshoot. The unhealthy pattern is pauses with an all-zero CNP column.
A one-second delta window hides microbursts. A fabric can pause for 200 µs every 10 ms and look calm at 1 Hz. If the numbers are too clean for the symptom you're chasing, sample at 100 ms before concluding anything.
Step 3 — correlate pause duration with queue depth and flow completion
Pull switch-side queue depth at the same cadence as the NIC-side counters and line three things up: pause rate, queue occupancy, and flow completion time.
| Time | rx_prio3_pause (+/s) | np_cnp_sent (+/s) | Egress queue depth | FCT p99 |
|---|---|---|---|---|
| 14:02:10 | 0 | 180 | 42 KB | 1.9 ms |
| 14:02:11 | 482 | 298 | 310 KB | 3.1 ms |
| 14:02:12 | 611 | 305 | 380 KB | 8.4 ms |
| 14:02:13 | 0 | 190 | 55 KB | 2.2 ms |
These numbers are synthetic and chosen to show shape. What you're looking for is the ratio. If queue depth peaks below the Xoff watermark while pauses still fire, the configured thresholds don't match what the hardware actually does — common when a switch uses a shared buffer and computes the per-priority watermark dynamically from occupancy. If FCT p99 spikes several multiples while aggregate bandwidth is flat, that's head-of-line blocking, not a capacity problem.
The Tuning Trade-offs Behind Most PFC and ECN Incidents
Head-of-line blocking and victim flows
A PFC pause is scoped to a priority, not a flow. When the queue in front of one congested egress port fills, every flow mapped to that priority stops — including flows bound for entirely uncongested ports. Those are victim flows, and in AI training they show up as a handful of ranks sitting idle while a collective waits on one straggler.
Two things reduce this in practice: keep the RoCE priority narrow (lossless traffic only), and put the congestion point as close to the sender as possible so the pause is short. Neither eliminates it. PFC doesn't know about flows and never will.
PFC watchdog, storm suppression, and the deadlock risk
PFC deadlock is the real hazard. If buffers fill with packets that can't drain because of a circular dependency between ports, the fabric stops permanently. PFC watchdog exists to break that.
When a port stays paused past a configured interval, the watchdog acts: on NX-OS it can drop or err-disable, and Arista EOS offers similar drop/errdisable actions. Every one of those actions deliberately breaks losslessness on that link. For RoCEv2 RC traffic that means drops, and RoCE RC handles drops with Go-Back-N behavior, which can retransmit far more than the single lost packet. A watchdog event is a trade: an unbounded stall for a bounded, expensive recovery.
Set the watchdog interval so it fires on genuine deadlock and not on a legitimate multi-millisecond tail. On a busy 800G fabric that line is narrower than most defaults assume.
ECN threshold and marking-rate choices
The ECN minimum threshold should be low enough that marking starts well before the queue approaches Xoff, and high enough that you aren't generating a CNP storm at idle. On many fabrics that lands somewhere in the range of a few tens of KB of instantaneous queue depth for the lossless priority. The correct number depends on buffer size, port count, and burst characteristics, so it isn't something you copy out of a vendor cookbook.
Receiver-side and sender-side DCQCN parameters matter too. On mlx5, the knobs show up under /sys/class/net/<iface>/ecn/ on many kernels; exact file names change between driver versions, so list the directory instead of trusting a stale blog post. min_time_between_cnps on the receiver and the rate-reduction factors on the sender are the two that most often explain "ECN is enabled but nothing changes."
Change one threshold at a time and re-run the same microbenchmark. Threshold changes interact, and a fabric that improves on aggregate bandwidth can regress on FCT p99 for exactly the collective you care about.
Mitigations for PFC Pauses and What They Cost
| Change | Effect | Cost |
|---|---|---|
| Order thresholds ECN min < max < PFC Xoff | PFC becomes a rare backstop | Requires knowing real buffer sizing |
Verify ECT codepoint end to end (cma_roce_tos) | Enables marking at all | None beyond the check |
| Narrow the lossless priority | Fewer victim flows | Less traffic gets lossless treatment |
| Enable PFC watchdog with drop action | Prevents permanent deadlock | Intentional loss, Go-Back-N retransmits |
| Raise ECN min threshold | Fewer CNPs, less rate collapse | Less headroom before PFC fires |
| Add headroom/buffer per port | Absorbs pause reaction window | Silicon and cost |
| Move congestion closer to senders | Shorter pauses, less victim impact | Topology constraints |
The ordering row is free. Do it first.
What to Watch For and What Remains Unconfirmed
Confirmed (from the specs and published work cited below): PFC's 802.1Qbb framing, per-priority watermark semantics, and hop-local scope; ECN's codepoints per RFC 3168; CNP-driven sender rate reduction per the DCQCN paper.
Inference from my own reasoning, not from a cited source: the ~500 KB per 400G port headroom figure is an order-of-magnitude estimate from bandwidth times reaction window, not a measured value. Treat it as a sizing sanity check.
Unconfirmed: I was given the headline and snippet of the New Electronics piece on Ethernet's reversal in AI clusters, not its full text, so I can't say what it specifically argues or which numbers it cites. Everything technical above stands on the specs and the DCQCN paper, not on that article. If the piece makes specific market-share or deployment claims, those need to be read directly.
What would change this post's advice: a fabric shipping with genuinely per-flow pause semantics would collapse the head-of-line blocking section. As of the specs I can verify, no such mechanism is deployed in RoCEv2 Ethernet.
Further Reading
- IEEE 802.1Qbb — Priority-based Flow Control — the amendment itself, from the IEEE 802.1 DCB task group.
- RFC 3168 — The Addition of ECN to IP — codepoints, marking semantics, and the CE/ECT rules.
- DCQCN: Congestion Control for Large-Scale RDMA Deployments (SIGCOMM 2015) — the sender/receiver algorithm behind most RoCEv2 congestion control.
- Ultra Ethernet Consortium — specification work for the next generation of Ethernet-based AI fabrics.
- SONiC — open-source NOS; useful for checking
show pfc countersand queue telemetry semantics against source. - rdma-core — home of
cma_roce_tosand the userspace RoCE tooling referenced above.


