Methodology
Every claim in every gem carries one chip. Here’s what each one means in plain English — and why we built the system the way we did.
A 12-person pilot, a single mouse experiment, and a 90,000-person meta-analysis all show up the same way in your inbox: as “studies.” That’s how the noise gets in. Before any claim leaves a Distilled Gems edition, it gets one of four chips — Strong, Moderate, Emerging, or Mechanistic-only. The chip tells you, at a glance, how confident the science actually is. No chip means we won’t say it.
What each chip means in plain English — with one worked example each.
A clear physiological or statistical mechanism, multiple converging human studies — including at least one meta-analysis or foundational controlled human trial — and expert consensus across credible voices.
Worked exampleMorning sunlight within 30 minutes of waking improves sleep quality. Backed by Berson, Dunn & Takao (Science, 2002), multiple human RCTs on phase-response curves, and consensus across Walker, Panda, Czeisler, and Foster.
A strong mechanism is established, and at least one high-quality human RCT or large prospective cohort confirms the effect — but no full meta-analysis yet, or some well-designed studies disagree on effect size.
Worked exampleEight thousand steps a day captures most of the longevity benefit of higher counts. Paluch et al., The Lancet Public Health, 2022 — large prospective cohort, replication still in progress.
Two or three solid human studies support the effect, no meta-analysis yet, mechanism is plausible but not fully mapped. Worth testing — still being refined.
Worked exampleResonance-frequency breathing at roughly six breaths per minute raises heart rate variability. Several small human studies, biologically plausible, no meta-analysis yet.
The biological, physiological, or behavioral mechanism is documented in animal models, lab studies, or small human pilots — but the clinical effect in healthy adults has not been demonstrated at scale. Use with appropriate caveat.
Worked exampleCompound X extends lifespan in mice by 30%. Robust mechanism in animals, intriguing — but no evidence yet that it does the same thing in humans living human lives.
Three checks. Every claim passes them in this order before it gets a chip.
What kind of study is this? A meta-analysis of RCTs sits at the top. A single rodent experiment sits at the bottom. Every source is mapped against the upstream evidence hierarchies.
Do other credible voices and other study designs land in the same place? Three independent RCTs plus a mechanism plus expert consensus is the bar for Strong. Disagreement drops the chip a tier.
If we don’t yet know — we say so. Every gem on a contested protocol carries a “what we don’t know yet” line. Replication uncertainty gets named, not buried.
No chip means no claim. If we can’t grade it, we won’t tell you to do it.
Distilled Gems didn’t invent evidence grading. The medical and research community has been refining it for thirty years. Our four tiers are a plain-English compression of the four most-used systems — translated for a reader who has five minutes on a Sunday morning, not five hours in a journal.
Those four systems lean on the same simple idea: some study designs are more reliable than others. Here are the seven main types in plain English, strongest at the top. Next to each is the chip it usually earns in a gem.
Reviews that combine many studies
StrongThis is a study of studies. Researchers gather every solid study on one question, then pool the results into a single big-picture answer. Combining thousands of people washes out the flukes of any one study. This is the gold standard.
The fair, randomized test
StrongPeople are split into groups by chance. One group gets the thing being tested. The other gets a dummy pill or the usual care. Because chance sets the groups, they start out even, so a difference at the end points to the treatment, not luck. It is the strongest kind of single human study. One small trial on its own usually earns a Moderate chip.
Studies that follow a group forward
ModerateResearchers track a large group of people for months or years and record what happens. They watch. They do not assign anything. This is how we learn the long-term effect of a habit we cannot test on purpose, like how you sleep or eat. It is weaker than a trial, because other parts of life can creep in and muddy the result.
Studies that look backward
EmergingStart with people who already have a condition, plus a similar group who do not. Then look back to find what was different between them. This is faster and cheaper than following people forward. It is also less reliable, because old memories and records have gaps.
Snapshots and surveys
EmergingMeasure a group of people once, at a single moment. These are good for spotting a link, like “people who sit more tend to weigh more.” But a snapshot cannot tell you which thing caused the other, or whether a third thing caused both.
A single story, or a handful
EmergingA close write-up of one person, or a few. They are useful for flagging something new or rare. But with no comparison group, you cannot tell if the result came from the treatment or from chance.
Lab and animal studies
MechanisticDone in petri dishes, in cells, or in animals like mice. They show that something is possible and hint at how it might work. But humans are not mice. A promising result here is a starting point, not proof in people.
This ladder is a guide, not a law. A large, careful study can outrank a small, sloppy one a rung above it. That is why we never lean on a single study. We look at where many studies land together, the convergence test from the steps above.
A high evidence grade is not a license to give medical advice. These four lines are non-negotiable.
Educational content only — never medical or financial advice.
When you open a Sunday gem and see a teal Strong chip next to a claim, you can act on it this week with high confidence the science is settled. An amber Moderate chip means the effect is real and worth doing, with a little more replication still to come. A rust Emerging chip means early but promising — worth testing, with the open questions named in plain sight; if it also carries a ↑ Rising tag, the supporting research is recent and building. A graphite Mechanistic-only chip means we’re showing it to you because the mechanism is interesting, not because the human evidence is in. Either way, the chip tells you the truth before the headline does.
The journals do the proving. We do the translating.