AI Conclusions

One click from results
to a written verdict

The AI agent reads the hypothesis, all statistical results, confidence intervals, and effect sizes — then produces a structured plain-English conclusion. Ship, reject, or investigate further.

Try free
What the agent sees

Full experiment context goes in — conclusion comes out

The agent doesn't just see numbers. It receives the full context — hypothesis, planned vs actual sample sizes, per-metric results with p-values, CIs, and deltas.

Hypothesis & context

Free delivery under €15 — checkout

Completed

Adding free delivery for orders under €15 will reduce friction and increase purchase conversion rate.

Key metric

Purchase CR

MDE

+1.5% relative

Sample sizes — planned vs actual

Control
planned 12,000 actual 12,500
Test 1
planned 12,000 actual 12,617

Statistical results

Metric Delta p-value 95% CI
Purchase CR key +0.5% 0.42 −0.3% / +1.3%
AOV −3.2% 0.005 −5.1% / −1.3%
GMV/user −1.8% 0.07 −3.7% / +0.1%

AI Conclusion

Generated in 4.2s · Review before publishing

✗ Do not ship

Overall conclusion

Free delivery did not increase purchase conversion but significantly reduced average order value. The hypothesis was not confirmed — instead of CR growth, we got deteriorating unit economics. No SRM detected, sample sufficiency confirmed on both groups.

Metric results

  • ·Purchase CR +0.5%, p=0.42 — within noise range, no reliable effect on conversion.
  • !AOV −3.2%, p=0.005 (95% CI: −5.1% to −1.3%) — significant drop. The most actionable finding.
  • ·GMV/user −1.8%, p=0.07 — not significant, but trend is consistently negative.
  • Insight: free delivery likely attracts price-sensitive users or shifts mix to lower-value items.

Summary

  • Decision:Do not ship. AOV declined significantly — worsening unit economics at scale.
  • Risks:Estimated loss of €0.13–0.51 per order from AOV compression.
  • Next test:Conditional free delivery from €40 — may increase AOV instead of reducing it.
Approach

Effect size over p-value theatre

Most tools return "significant / not significant". The agent reasons about what the numbers actually mean.

Effect size is the headline

A p=0.001 with a 0.02% delta is noise at scale. A p=0.07 with a −3.2% AOV drop is a red flag. The agent says this clearly instead of just "significant".

Honest about noise

p > 0.3 is labelled noise. p between 0.05 and 0.15 gets a nuanced read — "promising but inconclusive" — not a blanket "not significant".

Guardrail awareness

If a guardrail metric moves negatively — even when the key metric wins — the agent surfaces it and weights it in the decision. No silent regressions.

Available on all plans, including free Solo

Register and run your first experiment analysis — AI conclusions are included from day one.

Register free →