Analysis & Research for Experimentation

Many digital marketing teams waste time on A/B tests based on guesses rather than informed hypotheses. Effective testing relies on comprehensive user research, which identifies problems, prioritises them, and validates solutions. Multiple data sources enhance confidence in findings, transforming insights into actionable experiments.

Why Most A/B Tests Are a Waste of Time (And What to Do Instead)

There is a pattern that repeats itself endlessly in digital marketing teams. Someone runs a hunch. They change the colour of a button. They wait four weeks. Nothing moves. They try a new headline. Nothing moves. They start to wonder whether experimentation even works.

It does work. But not like that.

The teams who consistently win with A/B testing share one habit: they do not test guesses. They test answers to documented questions. Before anyone opens a testing tool, they have spent real time understanding what is actually happening with real users on their real website. That process is called user research, and it is the foundation that everything else rests on.

The Problem with Research That Goes Nowhere

User research is not a new idea. Most organisations do some version of it. The problem is not a lack of research it is a lack of connection between research and action. Insights get collected, reports get written, decks get presented, and then the document sits in a shared drive for eight months until someone is trying to justify a budget and digs it out.

This is called the implementation crisis, and it is the most expensive problem in conversion rate optimisation. The fix is not better reports. It is building research practices where every method leads directly to a test idea, a specific hypothesis, and a decision about what to do next.

Three Questions That Frame Everything

Before choosing any research method, you need to know what question you are trying to answer. There are really only three:

The first is: what problems exist? This is the exploration phase. You are casting a wide net. You do not know yet where the biggest issues are, so you are looking everywhere, reviewing the site against usability standards, watching session recordings, reading what users say when you ask them open questions.

The second is: which problems matter most? This is the focus phase. You now have a list of potential issues. Analytics tells you which pages see the most drop-off. Usability studies confirm whether an issue is widespread or just affected one person. Surveys help you understand how users talk about the problem in their own words.

The third is: did our solution work? This is validation. After a test concludes, research does not stop. You watch recordings of the winning variant. You check whether downstream behaviour changed, not just the primary metric.

Most teams only do fragments of this. They run analytics when something looks wrong, send an annual customer survey, and call it research. The teams who outperform them run all three phases as a continuous loop.

Why You Need More Than One Data Source

Any single research method can lie to you.

Analytics shows you a 74% drop-off rate on your pricing page. That is alarming, but analytics cannot tell you why. Are users leaving because they are confused about pricing? Because they need to compare with a competitor? Because the page loads slowly on mobile? The number alone tells you nothing actionable.

A session recording shows a single user scrolling back and forth between the features section and the call to action, looking lost. That is interesting, but it could be an anomaly. One person behaving strangely is not evidence of a systemic problem.

An on-site poll on that same pricing page collects 200 responses. Thirty-eight percent of them say some version of “I don’t know what’s included.” Now you have something. Three completely independent data sources quantitative analytics, qualitative behaviour observation, and direct user feedback, all pointing at the same underlying problem. That convergence is what elevates an observation from “something to investigate” to “something to act on.”

The research teams who get results run multiple methods simultaneously and wait to see where they overlap. The overlapping areas are where confident, high-return experiments live.

The Seven Methods

There is a collection of seven research methods that, when run together on a single project, generate the kind of overlapping evidence described above.

Heuristic review is the starting point. A trained evaluator, or better, a small team, walks through the website and scores each section against five principles: how much unnecessary friction exists, how clear the messaging is, how relevant the content is to the people arriving on that page, how compelling the value proposition is, and what motivates users to act. This requires no user recruitment and can be completed in a day. It surfaces issues fast and creates a shared language for the team.

Customer surveys are the most direct route to understanding how users think and feel. The most valuable question you can ask recent buyers is some version of: what almost stopped you from buying? The answers reveal objections and anxieties that the marketing team may never have considered, often expressed in language you can use directly in copy. The key discipline here is to send different surveys to different groups, loyal customers tell you what you are doing right, recent converters tell you what the journey was really like, and people who left without buying tell you what went wrong.

Usability studies put real users in front of the website and watch what they do. Not what they say they do, and not what they think they would do, what they actually do when asked to complete a realistic task. A critical insight from decades of usability research: you only need five participants per audience segment to surface approximately 85% of the significant usability problems. After five sessions, each new participant tends to surface issues you already documented. The method rewards intensity over volume.

On-site polls are short, one or two questions, shown directly on the website, in context, in the moment the user is making a decision. When someone is about to leave the checkout page, asking “What’s stopping you from completing your order today?” produces answers that a post-session survey cannot replicate, because the user does not have to reconstruct the experience from memory. One rule applies absolutely: never run polls on checkout pages. The cost of interrupting a user at the point of maximum intent is higher than any research value you will gain.

Analytics analysis is where most organisations start, and the single most common mistake they make is reading the aggregate numbers. A 2% conversion rate is meaningless on its own. Split by traffic source, by device, by new versus returning visitors, and suddenly a 2% average might be a 6% rate for email traffic and a 0.8% rate for paid social. That gap is where investigations begin. The discipline is to always ask “so what?” until an observation produces an action, not to stop at “the mobile drop-off is high” but to keep asking until you arrive at “we need to run session recordings filtered to mobile users who exit the pricing page.”

Session recordings and heatmaps are the most time-consuming method but provide the richest context. A heatmap compresses thousands of user journeys into a single visual layer, showing where users click, how far they scroll, where attention clusters. A session recording shows one individual journey in full. The practical approach is to begin with a broad pass of twenty to thirty recordings with no filtering, then return with filters based on what you noticed, rage clicks, specific exit pages, specific user segments. Filtering too early means missing the unexpected patterns that generate the most valuable insights.

Copy testing evaluates whether your words work. A headline that makes complete sense to the marketing team writing it may be entirely opaque to a first-time visitor. Copy testing tools recruit members of your target audience and show them a screenshot of your page, then ask structured questions about what they understood, what resonated, and what was unclear. A fifteen-person panel is enough to identify whether a clarity or relevance problem exists.

And a bonus

Data science and machine learning add a layer of predictive power that the seven methods above cannot reach on their own. Where analytics tells you what happened and research tells you why, machine learning can tell you what is likely to happen next and for whom. Clustering algorithms can segment your users into behavioural groups that no manual filter in GA4 would have surfaced, revealing that a specific type of visitor converts at three times the average rate and deserves its own targeted experiment. Predictive models can score individual sessions in real time, flagging users who are showing early signs of abandonment before they leave, and triggering personalised interventions.

Natural language processing can process thousands of open-ended survey responses or on-site poll answers in minutes, identifying themes and sentiment patterns that would take a researcher days to code manually. The practical entry point for most CRO teams is not building models from scratch, it is applying existing tools to the data you are already collecting: running clustering on your analytics exports, using NLP to accelerate survey coding, or feeding your A/B test results into a model that predicts which user segments are most likely to respond to a given type of change. Used this way, data science does not replace the research methods, it amplifies them.

Turning Research into Action

After running several of these methods, you will have a large collection of raw findings: verbatims from surveys, colour-coded notes from usability sessions, analytics screenshots, poll responses. The next job is to make sense of it.

The approach that works is to strip each method’s findings down to single observations, one idea per sticky note, and then cluster them visually. When clusters emerge, you name them as problem statements. A cluster supported by findings from four different methods is a high-confidence theme. A finding that only appeared in one source is a candidate for further investigation, not immediate action.

The synthesis step is best done as a workshop rather than a solo task. When the developers, designers, and product managers who will actually build the experiments help construct the research board, they have ownership over the conclusions. Teams that build the research together are far more likely to act on it.

Not every finding becomes an experiment. Some issues are clearly correct to fix without testing, a broken link, a missing trust badge, a typo. Fix those immediately. Some issues need more data before you can act, add the tracking first. Some issues are genuinely uncertain in their solution and have enough traffic to detect an effect, those become A/B tests. And some issues are interesting but exist on low-traffic pages where a test would take years to reach significance, those become hypotheses to revisit later.

The Hypothesis: Where Research Becomes a Test

An A/B test without a hypothesis is not an experiment. It is a gamble. The difference is in what comes before the change.

A proper hypothesis has four parts. It starts with the evidence, not “we think users are confused” but the specific finding: 38% of poll respondents said they did not know what was included in the Pro plan, and session recordings show users scrolling between the features section and the call to action without converting. That evidence is what separates a research-backed test from a hunch.

The evidence is followed by the proposed change, specific enough that any designer could build it without asking questions. Then the expected outcome: what user behaviour should change and in what direction. Then the measurement: which metric is primary, and what secondary metrics will be watched.

Without the evidence clause, the hypothesis is just an opinion. The research is what gives it weight.

Keeping Research Alive

One of the most persistent problems in research-driven teams is that insights from six months ago are invisible to the next project. A report was written, a presentation was given, and the findings effectively ceased to exist. New team members know nothing of them. The same usability study gets run again because no one knew the first one happened.

The solution is a living research database, a structured place where every insight is stored with its source, the page it affects, its current status, and whether it has been tested. When a new project starts, the first step is to search the database, not commission new research. Over time, the database also becomes evidence of the value of research itself: a record of which insights led to which experiments, and what the outcomes were.

What Changes When You Do This Well

The teams that build research-driven experimentation programmes stop arguing about which tests to run. The evidence makes the prioritisation obvious. They stop running tests that fail to move metrics, because their tests are answering real questions about documented problems. And they stop experiencing the implementation crisis, because the people who will build and analyse the tests were part of building the research that justified them.

The experiments stop being lottery tickets. They start being answers.

See you soon.

View Comments (7)

Leave a Reply

  1. […] It is tempting to treat documentation as admin, the chore after the real work. That is backwards. Documentation is what lets institutional knowledge outlast the people who leave, what lets a future test build on a past one, and what makes learning across many tests possible at all. Without it, there is no memory and no meta-analysis. […]

Subscribe to My Newsletter

Subscribe to my email newsletter to get the latest posts delivered right to your email. Pure inspiration, zero spam.

Discover more from Discuss Data Science, Machine Learning and Analytics

Subscribe now to keep reading and get access to the full archive.

Continue reading