Shelley Endicott
All work Test integrity, multiple brands and teams

When the test data can't tell the whole story

Several live A/B tests depended on a shared system that determined whether visitors saw the business-hours or after-hours experience. An outage—and later an outdated hours configuration—meant some visitors may have received the wrong experience. I traced the issue across the affected tests and made test-integrity checks part of our ongoing process.

My role
UX Researcher — identified and scoped test-integrity risk
Scope
Tests across four teams
When
June to July 2026
Tools
VWO, Confluence, Slack

1The setup

Many brand sites change what they show based on whether the call center is open. During business hours a visitor might see "Call to order." After hours they see a different button or a cart. Several A/B tests across brands were testing exactly that: phone versus cart, and different after-hours buttons.

All of those tests depended on one shared system knowing the call center's hours.

2What went wrong

Neither issue had been documented as a test-integrity risk. The affected tests continued running, and their results continued to be read without that context.

Source: both faults documented on the July 15 team meeting pages.

The test setup made the results harder to interpret. The brand teams ran these as holdbacks, with 99% of visitors in the new version and 1% held back as control. A control group that small already makes differences hard to read, and the faults added uncertainty about what visitors in either group actually saw.

Read limitation: split set by the brand teams. I read these tests as directional only.

3What I did

While prepping weekly meetings with all four teams, I noticed the same system showing up across different brands' tests. I traced which live tests depended on it, matched them against the two fault windows, and put the finding in front of every team the same week, with the specific tests at risk listed on each team's page.

Source: published to all four teams' meeting pages July 15, 2026.

Then I made it harder for this to slip by again:

One way to write splitsEvery page now reads the same: control is the small group, variant A the large one.
A weekly checkBefore results get discussed, every live test is checked for shared-system faults, build problems, and split issues.
Source labelsEach finding on a meeting page says where it came from, so teams can tell what's solid.
Backed by researchI tied the issue to published work on uneven experiment splits (Kohavi, Tang, and Xu, 2020).
A test result is only as good as the conditions it ran in.Nobody did anything wrong here. A shared system broke, and nothing connected that break to the tests depending on it. Part of my job is making sure teams know the conditions behind the evidence they act on.

Not recorded: how much the faults changed any test result, or whether teams re-ran or discounted the tests at risk.