How to Run Google Ads Experiments That Produce Trustworthy Results
Most Google Ads accounts run on accumulated folklore. Somebody changed the bidding strategy in March, performance moved, and the change became doctrine, even though a promotion launched the same week and the market moved with it. Before-and-after comparison is how most optimisation decisions get made, and it is the weakest form of evidence there is, because in a live auction environment everything else changes at the same time.
Google Ads ships a proper alternative: an experiments framework that splits traffic between your campaign and a modified copy, concurrently, in the same market conditions. Used with discipline, it converts opinions into evidence. Used casually, it manufactures confident noise. This article is about the discipline.
What the Tool Actually Does
Custom experiments let you create a trial version of a Search or Display campaign with one or more changes, then split traffic between original and trial by a percentage you choose. Both arms run simultaneously, which removes time-based confounds: seasonality, market shifts and promotions hit both arms equally. Setup takes minutes, you set the split and the dates, and the interface reports per-arm results with significance indicators. Demand Gen has its own A/B framework; Shopping campaigns are the notable gap, which matters for ecommerce accounts and pushes those tests toward geo-based designs instead.
Two mechanical choices shape result quality. Split traffic 50/50 unless you have a strong reason not to, because uneven splits extend the time to significance. And choose cookie-based splitting where offered, so the same user consistently sees one arm; search-based splitting reaches significance faster but lets users straddle arms, which muddies conversion attribution for longer-cycle purchases.
Testing Things Worth Testing
An experiment costs traffic, time and attention, so the hypothesis should be worth the price. The changes that justify experiments are structural: bidding strategy migrations (manual to Smart Bidding, CPA to value-based, target changes), the moves we described in value-based bidding; match type and keyword architecture changes; landing page swaps; RSA messaging strategies. Trivial changes (a new sitelink, one headline variant) rarely justify the framework, because the effect sizes are too small to detect in reasonable time on most accounts' volume. The brutal arithmetic: an account doing 200 conversions a month can reliably detect perhaps a 15 to 20 per cent difference within a month; detecting a 5 per cent difference needs volume most accounts do not have. Decide what effect size would change your decision, estimate whether your volume can detect it, and if it cannot, either run longer or accept that this question is not answerable by this method. The same statistics that govern A/B testing programmes on websites apply, and pretending otherwise is how accounts accumulate false learnings.
One hypothesis per experiment. A trial arm with new bidding and new copy produces a result you cannot act on, because you cannot tell which change did the work.
Running It Straight
The failure modes are behavioural, not technical. Set the duration in advance (four to six weeks for most accounts, longer for long conversion lags) and do not peek-and-stop: results that drift in and out of significance early are the statistical norm, and stopping at the first favourable reading is the most common way accounts fool themselves. Freeze both arms for the duration; a mid-test tweak to the control invalidates everything after it. Respect learning periods, because a trial arm with a new bidding strategy spends its first fortnight learning, and its early numbers are not its true performance; judge the back half of the window more heavily than the front. And mind conversion lag at the end: an experiment evaluated the day it finishes systematically undercounts the trial arm's most recent conversions.
When the result comes in, act symmetrically. A winning trial gets promoted through the built-in apply flow. A losing trial is a success too: it stopped you rolling a bad change across the account, which is the cheapest failure available. Record both in a testing register with the context, because a bidding result from peak season may not hold in February, and a register is what turns individual tests into institutional knowledge.
Beyond the Tool
Some questions outgrow campaign experiments. Account-level questions (does brand bidding pay, does the channel drive incremental revenue) need geographic holdout designs of the kind we covered in brand bidding and paid search incrementality, because in-platform splits cannot see cannibalisation of free channels. The principle is constant: concurrent control groups or it is not evidence.
An account that runs two or three honest experiments a quarter compounds an advantage that no amount of settings-tinkering matches, because its decisions stop being folklore. If you want help designing the testing roadmap or the statistical guardrails, our paid search team builds experiment programmes into every engagement. Get in touch.