Skip to article

DailyLens field note

How to Run an N-of-1 Experiment Without Fooling Yourself

Run a safer personal experiment with a baseline, one change, an anchored outcome, confounder notes, stopping rules, and cautious interpretation.

Adam Ciszewski9 min read

Share the Note

Share this field note with someone building a calmer system.

Article cover for How to Run an N-of-1 Experiment Without Fooling Yourself
Visual cover for How to Run an N-of-1 Experiment Without Fooling Yourself
notion image
On Sunday night, you move caffeine earlier, add a supplement, start a morning walk, and promise to close your laptop by 6 PM. By Thursday, your afternoons feel better. You want to keep whatever worked, but four variables changed at once. Memory supplies a confident explanation that the data cannot support.
The practical answer is to make the next test smaller. Write one question, observe a baseline, change one low-risk variable, measure one anchored outcome at the same moment, note a few plausible confounders, and decide the duration and stopping rules before you begin. Treat the result as a hypothesis until it repeats.
That process can improve exploratory self-tracking. It does not turn a personal log into a clinical trial or prove that an intervention caused the result.

First, use the term N-of-1 carefully

In clinical research, an N-of-1 trial has a specific meaning. The CENT reporting guideline describes a prospectively planned, multiple-crossover study in one participant, often using repeated treatment and control periods. A rigorous trial may include randomization, blinding, washout periods, clinical oversight, and a pre-specified analysis.
The AHRQ user's guide to N-of-1 trials covers design, ethics, statistics, implementation, and patient-clinician collaboration. That is a much higher bar than noticing that you slept better after a walk.
This article uses “personal experiment” for a lighter practice: structured, exploratory self-tracking of a reversible, low-risk behavior. Its purpose is to make a better personal decision, such as whether a meeting buffer is worth keeping. It is not research, diagnosis, or a safe route for testing medical treatment on yourself.
The distinction matters because clinical crossover designs work only under suitable conditions. The intervention must have a reasonably understood onset and offset, the underlying situation should be sufficiently stable, and effects from one period must not contaminate the next. The Cochrane Handbook's crossover guidance explains why carryover, period effects, and missing data can mislead even trained researchers.

Choose a question that can change a decision

“What gives me more energy?” is too broad. It invites you to collect everything and explain the pattern afterward.
A useful question names four things:
  • the behavior you will change;
  • the outcome you care about;
  • the moment you will measure it;
  • the decision the result could change.
For example:
On workdays, does a ten-minute walk without calls or audio immediately after my final meeting improve my capacity for normal evening demands at 6 PM enough to keep it as a default?
The intervention is specific. The outcome is functional, not a vague promise of “optimization.” The timing is fixed. The final decision is clear: keep the walk, revise it, or stop spending time on it.
If no plausible result would change what you do, the experiment is tracking for its own sake.

The One-Change Trial Card

Use this seven-part card before you collect the first test observation. Writing the plan in advance makes it harder to move the goalposts when an exciting result appears.

1. Decision

Write the decision in one sentence.
Example: “I am deciding whether to reserve ten minutes after my final meeting for a quiet walk on most workdays.”
This keeps the project proportional. You are deciding whether a small routine deserves a place in your day, not whether walking is universally effective.

2. Baseline

Observe what happens before changing the routine. Baseline data show the ordinary level, range, and direction of your outcome. They also expose whether Tuesday and Friday are fundamentally different kinds of days.
For a same-day behavior and a daily outcome, one full work week or seven comparable observations can be a practical planning window. Seven is a usability default, not a scientific threshold, and it does not make the later comparison statistically reliable. If your workdays vary sharply, the outcome is rare, or the intervention may act slowly, you need a longer or different design.
Keep the measurement process identical during baseline and test periods. If you create the scale after the intervention starts, your definition of improvement may already be influenced by what you hope to see.

3. One change

Change only the target behavior. Do not simultaneously adjust bedtime, caffeine, supplements, training, and meeting load.
Ordinary life will still change around you. The point is not laboratory control. The point is to avoid adding preventable ambiguity.
This principle is also the core of the DailyLens guide to supplement tracking for busy professionals, although supplement experiments require additional safety checks and often have uncertain onset or washout periods.

4. One primary outcome

Pick the result closest to the decision. For the walk example, use “post-work capacity” rather than five separate scores for mood, focus, stress, motivation, and sleep.
Define the scale with behavioral anchors:
  • 0: normal evening essentials feel beyond my current capacity;
  • 1: I can manage essentials after a recovery buffer;
  • 2: I can manage the normal routine, but nothing extra;
  • 3: I can be present and handle one optional demand;
  • 4: I have capacity for the evening I planned.
Record it at the same point after work. A 3 at 4 PM and a 3 at 8 PM do not describe the same decision context.
If you need help deciding what deserves measurement, start with personal analytics for energy, then reduce the list to one outcome for this experiment.

5. Confounder shortlist

Choose no more than three contextual factors that could plausibly move the outcome enough to alter your interpretation. For this example:
  • illness or acute pain;
  • an unusually short night of sleep;
  • a workday that ran more than two hours longer than planned.
Log them as short labels, not new dashboards. Confounders are context for interpretation. They are not permission to delete every inconvenient observation.
Predefine which days are genuinely non-comparable. If you exclude a day only after seeing its score, you are editing the story around the conclusion.

6. Duration and stopping rules

For a reversible behavior with an immediate or same-day expected effect, you might choose two working weeks of comparable test observations after the baseline. This is a planning convenience, not a clinical or statistical standard: ten or fourteen entries cannot convert an uncontrolled experiment into proof. Do not use this approach for interventions whose benefits, harms, withdrawal effects, or carryover may take longer.
Write stopping rules before you start. Stop if the activity causes pain, dizziness, marked distress, unsafe distraction, or any other concerning effect. Pause if illness, travel, or a major schedule change makes the planned comparison meaningless.
Clinical N-of-1 guidance asks researchers to pre-specify periods, outcomes, harms, and criteria for ending a trial. The behavioral SCRIBE single-case reporting guideline reinforces the value of documenting the design and planned analysis before results are known. Your private card is simpler, but the discipline is useful.

7. Interpretation rule

Define what would count as a meaningful result in terms of function, not only a decimal.
For example: “I will keep the walk if the typical test-day rating moves by at least one anchor, the improvement appears across ordinary workdays rather than one exceptional week, and the routine is easy enough to maintain.”
Also define the other outcomes:
  • No clear change: stop or redesign the experiment.
  • Mixed change: keep no conclusion; repeat under more comparable conditions if the decision matters.
  • Worse or adverse: stop and do not explain the harm away.
  • Too much missing data: the experiment is inconclusive.
An honest “I do not know” is a successful result if it prevents a false belief from becoming a permanent routine.

Read the pattern without manufacturing certainty

At the end, plot or list every observation in time order. Do not compare only the best baseline day with the best test day.
Ask:
  1. Did the level shift after the change?
  1. Was the direction already improving during baseline?
  1. Did the result appear on most comparable days or only a few?
  1. Were missed entries concentrated on difficult days?
  1. Did adherence fade as novelty wore off?
  1. Did the change improve the outcome enough to justify its cost?
Starting an experiment after an unusually bad week creates a classic trap: some improvement may have happened anyway as the extreme period passed. Expectations can change how you rate a subjective outcome. A new routine can also create short-lived enthusiasm. None of those possibilities makes the data useless. They make a causal claim too strong.
If the result matters and the behavior is safe to start and stop, you can repeat the comparison. A repeat that produces a similar functional change is more useful than one dramatic first run. Randomized periods or blinded comparisons can further reduce bias, but they move the project closer to formal single-case research and may require methodological and clinical support. The NIH Office of Disease Prevention overview describes the stronger standard expected of within-person randomized trials.

Medical and supplement boundaries

Do not use this framework to start, stop, switch, or change the dose of a prescribed medicine. Do not test an unapproved substance, combine supplements to “see what happens,” repeat an exposure that caused an adverse reaction, or use self-tracking to delay care for persistent, new, severe, or worsening symptoms.
Supplements can interact with medicines and may create meaningful risks. The FDA advises discussing supplements with a health professional, and the NIH Office of Dietary Supplements warns that risk can rise with high doses, multiple products, medication interactions, pregnancy, surgery, and some health conditions.
If the question concerns treatment, symptoms, pregnancy, medication, or a diagnosed condition, bring the question and your observations to a qualified clinician. A properly designed clinical N-of-1 trial is a collaboration, not a solo biohacking project.

Build your card today

Choose one reversible behavior you already consider safe, such as a five-minute meeting buffer, a fixed notification-off block, or a short no-input transition after work.
Write seven lines:
  1. Decision
  1. Baseline observations
  1. One change
  1. Primary outcome and anchors
  1. Three confounders at most
  1. Minimum duration and stopping rules
  1. Interpretation rule
Then collect the baseline before improving anything.
The discipline is intentionally modest. A personal experiment should leave you with a clearer next decision, not a stronger story than your evidence can carry.

Sources and further reading

If you want one place to record the change, outcome, and the context around it, explore the DailyLens biohacking and personal-pattern system.
Portrait of Adam Ciszewski

About the author

Adam Ciszewski

As a software engineer, tech team leader, and founder of DailyLens, he has spent years exploring cognitive optimization, biohacking, and physical recovery through supplementation and strength training. His work focuses on practical systems that help professionals manage energy, improve sleep, and develop healthier habits.

Share this article

Know someone who would find it useful? Send it their way.

From insight to practice

Build a System You Can Actually Return To.

DailyLens brings journaling, routines and personal signals into one calm workspace, so useful ideas become repeatable action.