The life of an affiliate is full of A/B tests. Either you need to split the creatives and understand who has the better CTR, then you want to cram an application collection form into the hell and understand whether this will make conversion better, then they brought new landings to affiliate network, and they seem to be purely visually normal, but will there be leads?
So what happens next? An arbitrator takes a couple of creatives, pushes them onto Facebook and, let’s say, gets the following results after delivery:
Kreo1: 2800 impressions — 100 clicks = 3.6 CTR
Kreo2: 3000 impressions — 100 clicks = 3.3 CTR
Kreo3: 3700 impressions — 100 clicks = 2.7 CTR
“Everything is clear as day!” our referee shouts in delight: “Creo1 has the biggest CTR, fuck everyone else!”
Or, let’s say, he has a couple of curses. And he dumps 100 clicks for each and looks like this:
The first prokla had 25% penetration and the second 37%.
“Ahaaa,” the arbitrator yells: “The first one is shit!”
And everything would be fine, but for some reason, after setup is assembled from all the elements tested in this way — it is not conversionit?♂️
Well, or conversionit, but the final values do not correspond to the tested ones. “Arbitration is one big random thing,” the arbitrator decides and goes to the plant.
And while he is walking, we will see where he was wrong, and for this we would like to plunge into the theory of probability and statistics, but we will not do this, because it is boring, tedious and abstruse, but we need to pour?
Those interestedsend it to Wikipedia and follow the links,For now, let’s understand one simple thing: the data that we received after the test may not be enough to make an unambiguous conclusion: is value A better than value B.
So how can we understand whether we have lost enough traffic or not? For such cases, smart people have long invented online statistical significance calculators, and you and Ilet’s deal with one of them.
Let’s take these, go to the “Testing Results” tab, enter the data from the first example with creatives there and see the result:

If you look closely, there is a slider at the bottom, which by default is set to 95%, which means that there is only a 5% chance that the CTR indicators of the creo will differ with further draining!

The same thing will happen with the second example, you can check it yourself.
So what sample size would need to be for the difference to be significant?
Let’s take our second example with procluses. Let me remind you that one got 25% and the other got 37% per 100 clicks. The difference between penetrations is 12%. We go to the first tab “Sample Size” and set 25% and 12% there.

We see that the sample size should be 2 times larger! We check, go back to “Testing results” and stupidly multiply our indicators by 2 (i.e. imagine that after draining another 100 clicks, the breakdown will remain the same, which, of course, is just an assumption that needs to be verified by tests).

Only now, with a 95% probability, we can be sure that we have tested the procluses.
This concludes the brief introduction to statistics, let’s count everything and turn it into a plus, gentlemen!

