Illustrative. The asymmetry is the argument: one calendar produces a story, the other produces four facts and one revert.
This is the cheapest habit in the entire playbook and the one most stores never adopt: change one thing, write down the date, wait, and compare like with like. Ship five changes in a week and revenue will move, but the movement belongs to all five, which means you cannot repeat it, cannot revert the harmful one, and have bought a result instead of a finding. A dated change log costs ten seconds per entry and is the difference between a store that compounds and one that thrashes.
The habit three chapters have already promised
Chapter 3.1 said a written change log turns step six of the 30-minute diagnostic from an hour of remembering into thirty seconds. Chapter 4.11 said sort order is unusually easy to change and unusually hard to attribute. Chapter 3.7 ended by insisting on one fix at a time.
All three were pointing here, and the underlying idea is one sentence: a result you cannot attribute is a result you cannot repeat.
Five changes in one week is not five experiments. It is one anecdote with five possible explanations. The Dropshipping Playbook
Why simultaneity is so expensive
"You cannot tell which one worked" undersells what is actually lost.
| What you lose | Consequence |
|---|---|
| The ability to repeat it | Next month you cannot do it again on another product |
| The ability to revert | If one change was harmful, its damage is hidden inside a net positive |
| The negative findings | A change that lost money looks like a smaller win, so you keep it |
| The transferable lesson | Nothing you learned applies to the next store or the next season |
| The ability to diagnose later | When revenue moves in three months, week 12 is a blur |
The third row is the quietly worst one. A batch of five changes where four help and one hurts nets out positive, so the batch is kept, and the harmful change stays in your store permanently, doing damage nobody will ever look for because the week it shipped was a good week.
The log itself
It does not need a tool. A spreadsheet, a text file, or a note in whatever you already open every day. Five columns:
| Column | Example | Why it earns its place |
|---|---|---|
| Date | 2026-08-26 | The whole point. Everything else is comparison against it |
| What changed | Free-shipping threshold $50 to $65 | Specific enough to revert exactly |
| Where | Shipping settings, all regions | So a later reset is detectable |
| Expected effect | AOV up, cart abandonment flat | Written before the result. This is the important one |
| What happened | Filled in two weeks later | Turns a log into a record of judgement |
Column four is the one people skip and the one that does the work. Writing what you expect before you see the outcome is what prevents the universal habit of deciding, afterwards, that whatever happened is what you were going for. It also makes a surprise legible: an unexpected result is only visible as a surprise if the expectation was written down.
App auto-updates, theme updates and supplier price changes all move revenue and none of them are your decision. Chapter 3.1 lists them among the changes that never notify you. Add them to the log when you notice them, because in three months they will be indistinguishable from things you did on purpose.
Comparing honestly
The log tells you when. Reading the result correctly is a separate discipline, and there are three ways it usually goes wrong.
- This week versus last week. Mixes weekday and weekend patterns and produces noise that reads as signal. Compare the same weekdays across four weeks, per chapter 3.1.
- Judging before the sample exists. Under 500 sessions a month, a fortnight may not be enough for anything to be readable. That is a reason to be honest about uncertainty, not a reason to shorten the window.
- Judging on the wrong number. A bundle raises revenue and can lower contribution; a first-row product outsells its old self because it is in the first row. Judge on the metric the change was supposed to move, chosen in advance, per chapter 5.1.
Why not just A/B test
Because most dropshipping stores do not have the traffic, and a badly powered test is worse than no test: it produces a confident answer from a sample that could not support one.
A split test needs enough conversions in each arm to distinguish a real difference from chance. At a 1.4% conversion rate and a few hundred sessions a week, reaching that on a modest effect takes months, during which the store cannot change anything else. That is rarely the right trade at this scale.
Sequential comparison, honestly logged, with like-for-like weekday matching, is the realistic method. It is weaker evidence than a properly powered split test and it is much stronger evidence than what most stores currently have, which is a memory.
What this looks like in practice
- One change per week. If that feels slow, note that the alternative is a month with no findings in it.
- Write the row before you make the change, including what you expect.
- Do not touch it for two weeks, unless something is obviously broken.
- Compare same-weekday, four weeks, on the metric you nominated.
- Fill in what happened, including when it was nothing, which is a finding and the most commonly discarded one.
- Revert anything negative immediately, which is only possible because you changed one thing.
That cadence is the spine of chapter 5.3, which turns it into a repeatable half-hour rather than a resolution.
Common questions
How long should I wait between changes?
Isn't this too slow when I need results now?
What if two changes have to ship together?
Do I need an A/B testing tool?
One number pays for the rest
What a single order is worth decides what every other fix can afford. Chapter 4.1 is where that number moves.
Go to chapter 4.1