How to assess · For hiring teams
How to Assess A/B Testing Skills When Hiring
The test formats that actually work for A/B Testing, what a strong answer looks like, sample questions and a scoring rubric you can use as-is.
The short answer
Assess A/B Testing with a task, not a conversation: design a test from a proposal, interpret a flawed result, ai-scored assessment (e.g. cohesyve) or programme conversation. Score it against written criteria you fix before you see any submissions, and weight the criteria that the role actually depends on.
- Defines the hypothesis, primary metric and minimum detectable effect before the test starts
- Calculates the sample size and duration and does not stop early on a green result
- Checks randomisation actually worked: sample ratio, pre-period balance
- Chooses guardrail metrics and knows what would make them stop a winning test
Paste a job description; Cohesyve generates a role-specific assessment and rubric. Ten candidates free, no card.
A/B testing looks simple — split traffic, compare metrics — and most organisations do it badly. Tests stopped the moment they turn green, metrics chosen after the fact, variants that were never actually randomised, and a winner declared from a sample that could not have detected the effect. The skill is not running tests; it is designing them so the answer means something and interpreting them without fooling yourself. This page covers how to assess A/B testing for product analyst, growth and experimentation roles: design, power, randomisation, interpretation and the discipline to report honestly.
Why A/B Testing is worth testing
Bad experimentation is worse than no experimentation, because it produces confident wrong answers. A team that peeks and stops early ships changes that do nothing; a team that tests ten metrics and reports the one that moved rewrites the product on noise. Testing shows whether a candidate understands power, peeking and selection, and that determines whether the experimentation programme creates knowledge or theatre.
What strong A/B Testing looks like
- Defines the hypothesis, primary metric and minimum detectable effect before the test starts
- Calculates the sample size and duration and does not stop early on a green result
- Checks randomisation actually worked: sample ratio, pre-period balance
- Chooses guardrail metrics and knows what would make them stop a winning test
- Interprets results with uncertainty and does not cherry-pick secondary metrics
- Understands common pitfalls: novelty effects, interference, seasonality, multiple variants
- Reports null results as useful and keeps a record of what was learned
Ways to assess A/B Testing
Design a test from a proposal
Describe a product change and the current metrics. Ask for the hypothesis, primary and guardrail metrics, sample size reasoning, duration, and what could invalidate the result. Forty-five minutes.
Pros
Cons
Best for Any experimentation role.
Interpret a flawed result
Provide a test readout with a sample ratio mismatch, a test stopped at day three, and a secondary metric presented as the win. Ask what can be concluded.
Pros
Cons
Best for Mid and senior roles.
AI-scored assessment (e.g. Cohesyve)
Generate an experimentation scenario from the job description — a design, a readout critique, a programme question — with a rubric. Each candidate receives a different variant; reasoning is scored in writing.
Pros
Cons
Best for Screening a pool.
Programme conversation
Ask how they would set up experimentation for a team that has never done it: tooling, process, culture.
Pros
Cons
Best for Senior and lead roles.
Cohesyve
Run a A/B Testing assessment on your next opening
Cohesyve generates a unique A/B Testing task per candidate from your job description, with the scoring rubric attached. Questions are different for every applicant, so they cannot be shared or looked up.
What to test
Design
Whether the test can answer the question.
Validity
Whether the result can be trusted.
Interpretation
Whether conclusions match the evidence.
Programme
Whether experimentation creates knowledge.
Sample A/B Testing questions
Why should you not stop a test the moment it shows significance?
EntryLook for Peeking inflates false positives; fixed horizon or sequential methods.
What is a sample ratio mismatch, and what does it tell you?
MidLook for Split deviates from the intended ratio; randomisation or logging is broken; results are not trustworthy.
The primary metric is flat but a secondary metric is up 12%. What do you conclude?
MidLook for Likely noise from multiple comparisons; treat as a hypothesis for a new test; do not ship on it.
How do you choose the minimum detectable effect?
MidLook for The smallest change worth acting on, given cost; drives sample size; a trade-off with duration.
Two teams run tests on the same page simultaneously. What is the risk and how do you manage it?
SeniorLook for Interaction effects; mutual exclusion, layered experiments, or accept and test for interaction.
Red flags
- Stops tests when they turn green
- Chooses metrics after seeing results
- Has never checked a sample ratio
- Cannot explain power
- Reports only winning tests
Scoring rubric
| Criterion | Weight | What strong looks like |
|---|---|---|
| Design discipline | 30% | Hypothesis, metric, power and duration set up front. |
| Validity checks | 25% | Randomisation and stopping rules are verified. |
| Interpretation | 25% | Conclusions respect uncertainty and avoid cherry-picking. |
| Programme thinking | 20% | Experimentation produces durable learning. |
Mistakes hiring teams make
- Testing tool knowledge instead of design judgement
- Not including a peeking or cherry-picking scenario
- Accepting a design with no power reasoning
- Ignoring guardrails
- Assuming analytics experience includes experimentation discipline
Roles that need A/B Testing
Common questions
Do I need an experimentation platform to assess A/B testing?
No. Design and interpretation are the skills, and both can be tested on paper with a proposal and a readout.
What is the best single A/B testing question?
Show a flat primary metric and a big secondary win and ask what to do. Multiple-comparison awareness and discipline show immediately.
Should product managers be assessed on A/B testing?
A light version, yes: what makes a result trustworthy, and why stopping early is a problem. They make the decisions that experiments inform.
How long should an A/B testing assessment take?
Forty-five minutes for a design exercise; thirty for a readout critique.
Cohesyve · Skill assessments for hiring
Test A/B Testing before the first interview
Generate a role-specific A/B Testing assessment from your job description and see who can do the work before you spend interview time on them.
1,500+
assessments completed
50%
faster time-to-hire
90%
completion rate
5 min
from JD to assessment
No credit card · 10 free candidates · Plans sized to your hiring volume
From the blog