Artive Ventures Working Paper AVWP 2026/01

The four-gate method: testing commerce ideas before building them

Working paper. Not peer reviewed. Comments welcome.

Abstract

Most new commerce products fail for a reason that could have been found cheaply: nobody had the problem badly enough to change what they do. This paper sets out the method the Artive Ventures lab uses to find that out early. Every idea passes four gates (research, prototype, pilot and spin out), and at each gate it faces one small test with one number, decided in advance. We describe the parts of a test, the evidence each gate asks for, and the rules for stopping. The method borrows from lean start-up practice, online controlled experiments and the pre-registration of studies, and adapts them to South African merchants, where small samples, payday cycles and informal trade make naive testing misleading.

1. Introduction

A venture studio has one scarce resource: the time of the people who build. Spending it on an idea nobody needs is the most expensive mistake it can make, and the most common. The lean start-up literature made this point for software (Ries, 2011; Blank, 2013). Commerce adds its own difficulties. Merchants are busy, samples are small, sales swing with paydays and holidays, and much trade in South Africa happens in cash, in chats and in spaza shops, where there is little data to test against.

The four-gate method is our answer. It is deliberately simple, so that it is followed. This paper describes it in full, so that founders who work with us know how their idea will be judged, and so that others can use it or criticise it.

2. The parts of a test

Every experiment in the lab is written down in five parts before it starts:

  • The question: one problem, written as a question a merchant would ask. If it cannot be written plainly, it is not understood yet.
  • The bet: what we expect the answer to be, and why. Writing the bet down first is a form of pre-registration: it stops the goalposts moving once results arrive (Nosek et al., 2018).
  • The smallest test: the least we can build, borrow or fake to find out. A landing page, a spreadsheet, a chat, a prototype.
  • One number: the single measure that decides the test, such as orders, repeat use, time saved or rand earned. Other measures are recorded but do not decide.
  • The call: keep going, change course or stop, with the reason written down.

The one-number rule is the hardest to keep. With many measures, something almost always improves, and a team that wants an idea to work will find it. Choosing the number in advance, and sizing the test to detect a change in it, is standard practice in online experimentation (Kohavi, Tang and Xu, 2020); we apply it to tests that are often far smaller than a website’s traffic.

3. The four gates

An idea moves through four stages. At the end of each is a gate: the evidence it must show before more time and money go in.

StageGateTypical evidence
ResearchA real problem that real merchants feelInterviews with people who have the problem; evidence they already spend time or money working around it
PrototypePeople use it without being asked twiceRepeat use of a rough first version, unprompted
PilotIt pays: revenue, retention or cost savedReal customers, real money or real stock, and one number that shows it pays
Spin outA reason to exist on its ownA business model that works at scale, and people to run it full time

The gates are ordered by cost. Research is cheap, a pilot is not, and a spin out commits people for years. Most ideas should stop at the first gate, and that is the method working, not failing.

4. Adapting tests to South African commerce

Three features of the market change how tests must be run.

  • Payday cycles. Spending in South Africa rises around month end. A test that runs for two weeks can mistake the calendar for the effect. Where spending matters, tests cover at least one full payday cycle.
  • Small samples. Many merchants have too few customers for a statistically clean test in a short time. Where numbers are small, we prefer larger, clearer effects (a behaviour that changes or does not) to small percentage lifts, and we say so when a result is directional only.
  • Informal and chat-based trade. Where there is no till data, the number may be counted by hand: orders in a chat, restocking trips avoided, repeat orders. It is recorded the same way every time.

5. Stopping rules

A test stops when its planned length is reached, or earlier if continuing would harm the merchants or shoppers involved. It does not stop early because the number looks good; early peeking is a well-known source of false positives in experiments (Kohavi, Tang and Xu, 2020). When a test stops, the call is written down with its reason. A clear no is published as a result, like a yes.

6. Limitations

The method trades rigour for speed. Small tests can miss real effects, and one number can hide a cost that matters. The gates describe evidence, not certainty: passing one makes the next investment reasonable, not safe. We expect to revise the method as the lab runs more experiments, and will publish changes in later versions of this paper.

7. Conclusion

The four-gate method asks one thing of every idea: show, as cheaply as possible, that someone needs it. It keeps the lab honest by fixing the question, the bet and the number before the answer is known. It is the way we judge our own ideas and the ones founders bring us.

References

  1. Blank, S. (2013). Why the lean start-up changes everything. Harvard Business Review, 91(5), 63-72.
  2. Kohavi, R., Tang, D. and Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge: Cambridge University Press.
  3. Nosek, B. A., Ebersole, C. R., DeHaven, A. C. and Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600-2606.
  4. Ries, E. (2011). The Lean Startup. New York: Crown Business.

Got a commerce problem worth a company? We want to hear it.

A few lines is enough: the problem, who has it and why now.

Pitch us