Artive Ventures Working Paper AVWP 2026/01
The four-gate method: testing commerce ideas before building them
Working paper. Not peer reviewed. Comments welcome.
Abstract
Most new commerce products fail for a reason that could have been found cheaply: nobody had the problem badly enough to change what they do. This paper sets out the method the Artive Ventures lab uses to find that out early. Every idea passes four gates (research, prototype, pilot and spin out), and at each gate it faces one small test with one number, decided in advance. We describe the parts of a test, the evidence each gate asks for, and the rules for stopping. The method borrows from lean start-up practice, online controlled experiments and the pre-registration of studies, and adapts them to South African merchants, where small samples, payday cycles and informal trade make naive testing misleading.
1. Introduction
A venture studio has one scarce resource: the time of the people who build. Spending it on an idea nobody needs is the most expensive mistake it can make, and the most common. The lean start-up literature made this point for software (Ries, 2011; Blank, 2013). Commerce adds its own difficulties. Merchants are busy, samples are small, sales swing with paydays and holidays, and much trade in South Africa happens in cash, in chats and in spaza shops, where there is little data to test against.
The four-gate method is our answer. It is deliberately simple, so that it is followed. This paper describes it in full, so that founders who work with us know how their idea will be judged, and so that others can use it or criticise it.
2. The parts of a test
Every experiment in the lab is written down in five parts before it starts:
- The question: one problem, written as a question a merchant would ask. If it cannot be written plainly, it is not understood yet.
- The bet: what we expect the answer to be, and why. Writing the bet down first is a form of pre-registration: it stops the goalposts moving once results arrive (Nosek et al., 2018).
- The smallest test: the least we can build, borrow or fake to find out. A landing page, a spreadsheet, a chat, a prototype.
- One number: the single measure that decides the test, such as orders, repeat use, time saved or rand earned. Other measures are recorded but do not decide.
- The call: keep going, change course or stop, with the reason written down.
The one-number rule is the hardest to keep. With many measures, something almost always improves, and a team that wants an idea to work will find it. Choosing the number in advance, and sizing the test to detect a change in it, is standard practice in online experimentation (Kohavi, Tang and Xu, 2020); we apply it to tests that are often far smaller than a website’s traffic.
3. The four gates
An idea moves through four stages. At the end of each is a gate: the evidence it must show before more time and money go in.
| Stage | Gate | Typical evidence |
|---|---|---|
| Research | A real problem that real merchants feel | Interviews with people who have the problem; evidence they already spend time or money working around it |
| Prototype | People use it without being asked twice | Repeat use of a rough first version, unprompted |
| Pilot | It pays: revenue, retention or cost saved | Real customers, real money or real stock, and one number that shows it pays |
| Spin out | A reason to exist on its own | A business model that works at scale, and people to run it full time |
The gates are ordered by cost. Research is cheap, a pilot is not, and a spin out commits people for years. Most ideas should stop at the first gate, and that is the method working, not failing.
4. Adapting tests to South African commerce
Three features of the market change how tests must be run.
- Payday cycles. Spending in South Africa rises around month end. A test that runs for two weeks can mistake the calendar for the effect. Where spending matters, tests cover at least one full payday cycle.
- Small samples. Many merchants have too few customers for a statistically clean test in a short time. Where numbers are small, we prefer larger, clearer effects (a behaviour that changes or does not) to small percentage lifts, and we say so when a result is directional only.
- Informal and chat-based trade. Where there is no till data, the number may be counted by hand: orders in a chat, restocking trips avoided, repeat orders. It is recorded the same way every time.
5. Stopping rules
A test stops when its planned length is reached, or earlier if continuing would harm the merchants or shoppers involved. It does not stop early because the number looks good; early peeking is a well-known source of false positives in experiments (Kohavi, Tang and Xu, 2020). When a test stops, the call is written down with its reason. A clear no is published as a result, like a yes.
6. Limitations
The method trades rigour for speed. Small tests can miss real effects, and one number can hide a cost that matters. The gates describe evidence, not certainty: passing one makes the next investment reasonable, not safe. We expect to revise the method as the lab runs more experiments, and will publish changes in later versions of this paper.
7. Conclusion
The four-gate method asks one thing of every idea: show, as cheaply as possible, that someone needs it. It keeps the lab honest by fixing the question, the bet and the number before the answer is known. It is the way we judge our own ideas and the ones founders bring us.
References
- Blank, S. (2013). Why the lean start-up changes everything. Harvard Business Review, 91(5), 63-72.
- Kohavi, R., Tang, D. and Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge: Cambridge University Press.
- Nosek, B. A., Ebersole, C. R., DeHaven, A. C. and Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600-2606.
- Ries, E. (2011). The Lean Startup. New York: Crown Business.
