1The setup
Many brand sites change what they show based on whether the call center is open. During business hours a visitor might see "Call to order." After hours they see a different button or a cart. Several A/B tests across brands were testing exactly that: phone versus cart, and different after-hours buttons.
All of those tests depended on one shared system knowing the call center's hours.
2What went wrong
- June 26An outage told sites the call center was closed for about four hours during business hours. Visitors in that window may have been shown the after-hours experience during business hours.
- July 7The hours chart the system used was confirmed out of date. The after-hours window itself may have been wrong for an unknown stretch of time.
Neither issue had been documented as a test-integrity risk. The affected tests continued running, and their results continued to be read without that context.
Source: both faults documented on the July 15 team meeting pages.
The test setup made the results harder to interpret. The brand teams ran these as holdbacks, with 99% of visitors in the new version and 1% held back as control. A control group that small already makes differences hard to read, and the faults added uncertainty about what visitors in either group actually saw.
Read limitation: split set by the brand teams. I read these tests as directional only.
3What I did
While prepping weekly meetings with all four teams, I noticed the same system showing up across different brands' tests. I traced which live tests depended on it, matched them against the two fault windows, and put the finding in front of every team the same week, with the specific tests at risk listed on each team's page.
Source: published to all four teams' meeting pages July 15, 2026.
Then I made it harder for this to slip by again:
Not recorded: how much the faults changed any test result, or whether teams re-ran or discounted the tests at risk.