An experiment log records what you expected, what you changed, what you observed and what you decided. Its purpose is organisational learning, not turning every change into a success story. A null, mixed or unusable result is still worth recording when it prevents the same uncertain test from being repeated.
Write a falsifiable hypothesis
Use a structure such as: “For [defined audience or traffic segment], changing [one named element] from [current state] to [test state] may change [defined metric], because [evidence or reasoning].” Avoid hypotheses that contain several changes or outcomes. If any observation can be described as success, the hypothesis cannot guide a decision.
Record the baseline and decision rule
Link to the measurement definition. Record the baseline period, data source, filters, event definition and known data gaps. Before starting, state what pattern would support keeping, revising or reverting the change. A decision rule need not pretend statistical certainty; it must make the judgement criteria visible.
| Before the change | After the observation |
|---|---|
| Question and hypothesis | Observed metric and period |
| Baseline and measurement method | Difference from baseline or comparison |
| Exact change and scope | Instrumentation or delivery issues |
| Start, stop and rollback conditions | Interpretation and alternative explanations |
| Decision owner | Keep, revise, revert or collect more evidence |
Change one interpretable thing
Record the exact pages, message, audience rule, distribution route or interface element changed. Save the prior version and note any simultaneous releases. If several elements must change together, describe the test as a package; do not later claim that one component caused the result.
Protect the audience
An experiment label does not remove ordinary obligations. Do not knowingly degrade accessibility, conceal material information, manipulate consent or expose personal data. Name stop conditions for broken forms, severe task failure, unexpected data collection or other material harm. Verify the rollback route before starting.
Observe without rewriting the history
At the end of the planned window, copy the results into the log with the extraction date. Preserve the original hypothesis and decision rule. Record instrumentation failures, low volume, unusual events and any changes made during the period. If the measurement definition changed, do not merge the two periods as though they were comparable.
Use neutral language: “the observed rate was higher in this period” rather than “the message drove growth” unless the design and evidence support that causal conclusion. Small samples and noisy channels often justify a cautious decision or another observation window.
Make and record the decision
- Keep: retain the change and state why.
- Revise: create a new hypothesis rather than editing the old one.
- Revert: restore the prior version and record the reason.
- Inconclusive: state what evidence is missing and whether further observation is worthwhile.
Copy-and-fill log
Experiment ID; owner; decision supported; hypothesis; evidence; baseline; exact scope; metric and guardrail; start and stop; rollback; observed data; data-quality notes; interpretation; alternative explanations; decision; decision date; link to source exports and versions.
Use the content checklist on the test material and record its approved scope in the marketing brief.