A/B Testing
- In Turkish
- A/B Testi
In short
A/B testing is an experiment that shows two versions of a feature to random groups of real users and measures which one performs better on a chosen metric.
What is A/B testing?
A/B testing compares two versions of something, such as a page layout, a feature, an email subject line, or a recommendation algorithm, by showing each version to a different random group of real users. Version A is usually the current experience, called the control, and version B is the change, called the variant. The team then measures which version does better on a chosen metric, such as sign-up rate, checkout conversion, or time to complete a task.
Users are assigned to groups at random, typically by hashing a stable user ID so each person keeps seeing the same version. Before starting, the team picks one primary metric and calculates how many users it needs, then runs the experiment long enough to cover normal weekly patterns. At the end, a statistical test checks whether the difference is significant, meaning unlikely to be random noise. Feature flags are the usual mechanism for serving each version, and stopping early as soon as B looks ahead is a classic mistake that produces false winners.
It's like a bakery selling two recipes of the same cookie side by side for a month and counting which one customers buy more, instead of asking the staff which they prefer. A/B testing is common in product design, e-commerce, marketing, search ranking, and pricing. Variants include A/B/n tests, which compare more than two versions, and multivariate tests, which change several elements at once.
Unlike most testing terms, A/B testing doesn't check whether code is correct; it measures how users respond to a change that already works. It is also different from a canary deployment. A canary sends a small share of traffic to a new release to catch errors before a full rollout, while an A/B test deliberately splits users to compare business outcomes between two working versions.
Key takeaways
- A/B testing compares a control (A) and a variant (B) with real users.
- Users are randomly and consistently assigned to one version.
- Choose the metric and sample size before the experiment starts.
- Statistical significance separates real effects from random noise.
- It measures user behavior, not code correctness.
Example
import { createHash } from "node:crypto";
// Hash the experiment name and user ID so each user always gets the same group
function assignVariant(userId, experiment) {
const hash = createHash("sha256").update(`${experiment}:${userId}`).digest();
return hash[0] % 2 === 0 ? "A" : "B"; // roughly a 50/50 split
}
const variant = assignVariant("user-42", "checkout-button-text");
const buttonText = variant === "A" ? "Buy now" : "Complete purchase";
// Record which version the user saw, so conversions can be compared per group
analytics.track("experiment_viewed", { experiment: "checkout-button-text", variant });Readers ask
How long should an A/B test run?
Long enough to reach the sample size you calculated in advance, and usually at least one or two full weeks so that weekday and weekend behavior are both included. Ending a test the moment results look good often leads to false conclusions.
What is statistical significance in A/B testing?
It is a measure of how unlikely the observed difference would be if the two versions actually performed the same. Teams commonly require a significance level of 5 percent before declaring a winner.
What is the difference between A/B testing and a canary release?
A canary release exposes a new version to a small share of traffic to catch errors before rolling it out to everyone. An A/B test splits users between two working versions to learn which one produces better results.
See also
- Feature FlagDevOps & Cloud, p. 21A feature flag is a switch in code that turns a feature on or off at runtime, letting teams deploy code without releasing it to every user at once.
- Canary DeploymentDevOps & Cloud, p. 6A canary deployment releases a new software version to a small share of users first, checks its health, and then gradually rolls it out to everyone.
- MetricsDevOps & Cloud, p. 36Metrics are numeric measurements of a system collected over time, such as request rate, error rate and CPU usage, used for dashboards, alerts and planning.
- MVPTeams & Process, p. 13An MVP is the simplest version of a product that real users can use, built to test a key assumption and learn as much as possible with the least effort.
- HashingSecurity, p. 14Hashing is the process of turning any input into a fixed-length value with a one-way function, used to verify data integrity and store passwords safely.
Spotted a mistake or something missing on this page?Suggest an edit