Why we burn-in every switch for 48 hours
← All insights

Why we burn-in every switch for 48 hours

RELIABILITY 10 FEB 2026 7 MIN

Every OKTET switch spends 48 hours powered on, under load, at elevated temperature, before it is allowed near a shipping box. This is not a marketing line. It is a line item on our build cost, it slows down throughput in the factory, and we keep doing it because the alternative is worse.

The shape of the failure curve

Electronic hardware fails on a well-documented curve. Plot failure rate against time and you get a bathtub: a high rate early on, a long flat valley in the middle, and a slow rise at the end as components wear out. The early spike is called infant mortality. It is dominated by manufacturing defects — a marginal solder joint, a capacitor that was already out of spec when it was reeled, a power MOSFET with a microscopic die crack that will propagate the first time it gets hot.

The important property of infant mortality is that it is front-loaded in time. A unit that survives its first few days of operation under stress has a dramatically lower chance of failing in the next several years. Burn-in is the deliberate decision to consume that high-risk window in our building, on our bench, instead of in a customer’s wiring closet three weeks after install.

What 48 hours actually looks like

A switch coming off the line is racked in a thermal chamber held at 50 to 55 degrees Celsius — well above a normal closet, below the rated maximum. We then drive it hard:

  • All ports up, negotiated at full speed, with traffic generators pushing line-rate frames in a mix of sizes including 64-byte minimums, which are the worst case for the switching ASIC’s packet-per-second budget.
  • PoE ports loaded against electronic loads that pull near the per-port and total budget ceiling, because the PoE subsystem and its DC-DC converters are among the hottest, hardest-working parts of the board.
  • Power cycled on a schedule. Thermal cycling — heating and cooling the board repeatedly — is far better at surfacing weak solder joints than holding a steady temperature, because it works the joints mechanically through expansion and contraction.

Throughout, the unit logs its own internal temperatures, fan RPM, PoE consumption, and any correctable or uncorrectable memory errors. We are not just waiting to see if it dies. We are watching the telemetry for drift.

Why 48 and not 24, not 96

The number is a deliberate trade, not a round figure. Accelerated-life models — the Arrhenius relationship for temperature-driven failure mechanisms — let you estimate how much field time a given temperature buys you per hour on the bench. At our chamber temperature, 48 hours of stress maps to a meaningful slice of the early-life window for the failure mechanisms we care about most.

Pushing to 96 hours yields diminishing returns: the marginal units that 48 hours would catch have largely already been caught, and you start spending bench time and energy to flush out failures that are statistically rare. Dropping to 24 hours measurably raises the escape rate of thermal-joint defects in our own historical data. So 48 is where the curve flattens for our boards. We revisit the number whenever the failure data tells us to.

What we do with a failure

A unit that fails burn-in is not quietly swapped and forgotten. It is logged, and if a particular board revision or component lot starts producing failures above its baseline rate, that is a signal that propagates back up the line. A bad capacitor lot gets quarantined. A reflow-profile problem gets corrected at the oven. Burn-in is a screen, but it is also a sensor on the health of the whole manufacturing process, and the second job is arguably the more valuable one.

What this does not catch

Burn-in is honest about its limits. It is excellent at infant mortality and at gross thermal and power defects. It does not catch wear-out failures decades away, it does not simulate every electrical transient a real building will throw at a unit, and it cannot find a firmware bug that only appears under a specific traffic pattern in production. Those belong to other programs — design margin, surge testing, firmware regression.

What burn-in buys is specific and measurable: the units that reach you have already survived the most failure-dense period of their lives. That is the difference between a switch you install and forget and one you end up driving back to the site to replace. We would rather absorb that cost on the bench, where it is cheap, than ship it to you, where it is not.