Test Size for New Gaming Format: Define primary outcome and eligible events for decision support; Estimate 5,000 visits from 500,000 impressions at 1% qualified rate; Use powered comparison with baseline, effect size, and significance for lift tests
Image: Gaming Ad Guide

Attribution

Part of Gaming campaign economics

Estimating a test size for a new gaming format

Estimate usable gaming test volume and cost from the decision, and distinguish feasibility arithmetic from a powered lift study.

Size a first test from the decision it must support. Define the primary outcome, estimate the eligible delivery and observable events needed, then cost that volume.

If the decision requires evidence of incremental lift, a delivery forecast alone cannot size the study: the comparison design needs its own sample calculation.

Choose the question and counting unit

A delivery test asks whether the new format can produce enough qualified, measurable impressions. A response test asks whether it produces enough useful actions to justify another buy.

A lift test asks whether it changes an outcome relative to a suitable comparison group. Each question needs different evidence.

Specify the games, devices, Australian eligibility rule, format, outcome window and unit of analysis. Impressions are not distinct devices, and repeated actions by one customer are not independent customers.

Estimate a feasible volume and cost

Use a proposal-specific supply forecast and a range of assumptions for the steps you need to observe. If no relevant historical response rate exists, label the rate as an assumption and show a conservative case. Do not borrow a rate from another format as though it were evidence for this one.

For hypothetical arithmetic only, suppose 500,000 eligible counted impressions produce qualified visits at an assumed 1%, and 2% of those visits produce a useful action. That gives 5,000 visits and 100 actions.

At an assumed AUD $20 CPM, media would cost AUD $10,000 before creative and measurement. These rates and prices are not gaming benchmarks. The calculation checks feasibility; it does not show that 100 actions are statistically sufficient or caused by the ads.

Check whether the test can decide

For an operating test, set minimum qualified delivery, the useful-action definition, a spend cap and a review date that allows the outcome window to mature. State in advance what leads to expand, revise or stop. Record material changes to creative or inventory.

For a powered comparison, specify the baseline outcome rate, smallest effect worth detecting, significance level, desired power, group allocation and unit of randomisation. A sample calculation must account for the actual design, exclusions and repeated users. NIST's sample-size example for a proportion illustrates the dependence on baseline, effect, significance and power; it is not a ready-made gaming lift calculation.

Lift studies compare a test group with a comparison group. A user-based Conversion Lift setup is one option where available; whether it can be used for the proposed gaming placement must be checked separately.

If a powered study is unaffordable or unavailable, narrow the question. A smaller test can establish whether the format runs, reporting works or the response path has an obvious problem. Describe that as a feasibility finding, not measured lift.

Use the actual media billing basis to cost the required volume, then add creative, verification and study costs. A test is large enough when the expected usable data can answer its stated question within the approved budget, not when it reaches an arbitrary impression count.

More from Attribution

Attribution

Evaluating a low-cost game placement with weak outcomes

Decide whether to repair, narrow or stop a low-cost game placement by checking total cost, eligible delivery and useful outcomes.