Article
Email A/B Testing Best Practices Explained
mastercard.com
Quoted on this wiki
Every place a page here uses this source, in the order the words come in it.
“It runs the test until one of the two variants hits a majority in terms of, say, open rate, and then it declares that email to be the winner of the A/B test.” The problem is a lack of statistical significance. You might be asking, “statistical what?”
And that’s why it’s so dangerous when A/B testing tools don’t take into account statistical significance. The “winning” email is the one with the higher open rate, and it will automatically get blasted out to the rest of your list. A more sophisticated test might reveal that the winning email actually performs far worse, but there’s no way to know using most of these systems. The software itself enables and rationalizes bad decision making. “Those clients have heard that email marketing is important and as a result, it’s pretty easy to sell them on an email testing product — most don’t even know that they should be looking for statistical significance.” Think of it this way: When you’re a marketer performing A/B tests, what you’ve been tasked with generating is insight. And any result, winner or loser, is still insightful. But an inconclusive result? Too many of those and you’ll stop performing A/B tests and, by extension, stop paying for that A/B testing add-on.
Simply put, statistical significance is the likelihood that an experiment is not due to random chance, entirely meaningless or just flat out wrong. “To achieve an acceptable level of statistical significance, a data set needs to be large enough, and generally the smaller the effect being measured, the larger the data set needs to be.” Statistical significance is measured in terms of % confidence or what’s called a p-value. So, let’s say you want to be 95% sure that the change you’re making to your emails will have a positive effect. Then, you’d want to see test results that show a 95% confidence level or a p-value of 5%.
The problem is a lack of statistical significance. You might be asking, “statistical what?” “Simply put, statistical significance is the likelihood that an experiment is not due to random chance, entirely meaningless or just flat out wrong.” To achieve an acceptable level of statistical significance, a data set needs to be large enough, and generally the smaller the effect being measured, the larger the data set needs to be.
“Sounds simple enough, but from a statistical point of view, this is a terrible idea.” The problem is a lack of statistical significance. You might be asking, “statistical what?”
Most marketers are impressed by (and sold on) a feature where you can run a test using a small sample of your full email list. It runs the test until one of the two variants hits a majority in terms of, say, open rate, and then it declares that email to be the winner of the A/B test. At this point, it auto-deploys the winning email to the rest of your list. Sounds simple enough, but from a statistical point of view, this is a terrible idea. “The problem is a lack of statistical significance.” Simply put, statistical significance is the likelihood that an experiment is not due to random chance, entirely meaningless or just flat out wrong.