Methodology
How the score is built.
Every intervention gets three 0–10 axes—Evidence, Benefit, and Safety— blended into a weighted composite, then mapped to a 14-grade letter scale. The same arithmetic is used across categories, but each score only applies to the population, use, outcomes, dose, and duration described in that entry.
The unit being scored
Evidence does not belong to a molecule in the abstract. It belongs to a claim: an intervention used by a defined population, for a defined purpose, compared with something else, over a stated period. Vitamin D for confirmed deficiency is not the same claim as vitamin D for preventing disease in an already-sufficient adult. The current scorer uses the first clause as the headline score, so that clause must name the exact claim and population it represents. Other contexts remain visible in the entry; mixed-use rows are being migrated toward separate context-specific claim blocks.
The formula
Composite = 0.45·Ev + 0.15·Bn + 0.40·Sf
(each axis 0–10 → composite 0–10)
Evidence carries the heaviest weight because the brand promise is evidence-first. Safety is next because a large effect with bad tolerability or serious downside should be pulled down visibly. Benefit still matters, but it cannot dominate the grade on its own. A numeric composite is descriptive, not a recommendation: an intervention that is effective and dangerous can still produce a mid-range number. Entries with a high-risk evidence profile are labeled so readers do not mistake the composite for a safety score; the Benefit and Safety axes remain separate and equally visible. Cost and legal access aren't scored at all: they vary too much by country, insurer, and supplier to belong in a universal grade. Each compound page still notes its legal status as context — it just doesn't move the letter grade.
Evidence (45%)
How much is actually known about this in humans?
- 9–10 — Very strong: Multiple high-quality RCTs, robust meta-analyses, or decades of human outcome data.
- 7–8 — Strong: At least one large, well-run RCT plus convergent secondary evidence.
- 5–6 — Moderate: One solid RCT, or several consistent observational studies.
- 3–4 — Weak / mixed: Small trials, conflicting results, or thin data.
- 0–2 — Preclinical or none: In-vitro, animal-only, anecdotal, or no usable data.
Study design alone does not settle certainty. Risk of bias, consistency, directness, precision, registration, attrition, outcome choice, and publication bias all matter. Conference abstracts, preprints, registry results, and company toplines are labeled provisional; they should not silently receive the same weight as a peer-reviewed trial or regulator-reviewed evidence package.
Benefit (15%)
If the claim is correct, how big is the effect?
- 9–10 — Very large: Mortality or disease-incidence shifts, or biomarker changes that translate clinically.
- 7–8 — Large: Reproducible effect sizes that show up in real-world outcomes.
- 5–6 — Moderate: Real but modest. Often only on surrogate biomarkers.
- 3–4 — Small: Detectable in trials, hard to feel in life.
- 0–2 — Negligible: Within noise or null on hard endpoints.
Safety (40%)
Inverted from risk: higher score = safer.
- 9–10 — Very low risk: No meaningful safety signal at doses represented by the scored evidence.
- 7–8 — Low: Generally well-tolerated. Minor side-effects in a minority.
- 5–6 — Moderate: Real side-effect profile, drug interactions, or monitoring requirements.
- 3–4 — High: Serious adverse events reported at therapeutic or studied doses; clinical use commonly includes monitoring.
- 0–2 — Severe: Fatal potential, a narrow therapeutic margin, or multiple severe harms.
What about cost and legality?
Neither feeds the score. How hard a compound is to obtain — prescription vs. over-the-counter, cash price, schedule status — depends on where you live, your insurer, and your supplier, so it can't sit inside a universal grade without distorting it. Each compound page records its legal status as a note for context; it just doesn't change the letter grade.
From composite to letter grade
The 0–10 composite gets mapped to a 14-grade letter scale. Most grades cover a 0.6-point band; the top four (A− through A+) are tightened to 0.3 points because the best-evidenced interventions cluster near the top of the scale. S covers everything from 8.5 up; F catches everything below 2.2.
Toxins use a separate scale
The toxin model (heavy metals, pollutants, and lifestyle exposures) is designed to characterize evidence of harm and exposure rather than an intervention-benefit axis. It reports exposure concern: harm magnitude, evidence of that harm, and exposure prevalence. A higher score describes a stronger and/or more widespread harm signal; it is not an instruction for an individual reader.
Exposure concern = 0.40·Magnitude + 0.30·Evidence + 0.30·Prevalence
Frontier potential — a second, separate number
The composite answers one question: how well established is this? On most entries that is the only question worth asking. On the frontier it is not.
An example from our own database. Epithalon, a pineal peptide resting largely on one research tradition, graded C+. Rapamycin, the most reproducibly life-extending compound in animal biology, graded C. The formula produced that ordering honestly: Safety is 40% of the composite, and Epithalon scores high on Safety because almost nobody has studied it. Rapamycin scores lower because it has been studied enough to document real immunosuppression. A high Safety score earned by absence of scrutiny outweighed one earned under it.
The composite was not wrong about what is established — rapamycin has no proven human longevity benefit. It was simply the only number on the page, so it got read as a measure of how interesting something is. Rather than reweight the formula and silently re-grade every entry, some entries now carry a second number alongside the grade.
What it is
A 0–10 judgment answering a different question: if this worked in humans, how much would it matter? Rapamycin scores 9 on that axis while grading C on evidence, and both numbers are true at once.
Three rules that keep it honest
- It never enters the composite. Nothing about potential can raise a grade. This is the whole design — a promise number folded into a headline score is the mechanism by which supplement marketing works, and keeping the two apart is what makes publishing the optimistic one defensible at all.
- Every score states what it rests on, and what would sink it. Each entry carries both the specific result behind the ceiling and the reason it may not be reached, shown side by side at equal weight. A ceiling without a stated basis is an opinion with a decimal point; a ceiling without its counterweight is marketing. An entry missing either fails our build.
- It is scored against published anchors, not vibes. 9–10 means a large effect on a hard outcome replicated across labs and species. 5–6 means a coherent mechanism with consistent cell and animal data. 1–2 means tradition or marketing with no mechanism to support it. 0 is reserved for tested in humans and failed, which is the value that separates disproven from untested — a distinction the evidence grade alone cannot express.
What it is not
It is editorial judgment, labelled as such wherever it appears, and it is the least defensible number on this site — unlike Evidence and Safety, it does not come from documents. It is a statement about an idea, not a prediction that any reader will benefit, not a recommendation, and not a substitute for the evidence grade sitting next to it. Coverage is deliberately partial: it appears only where we can write a real basis and a real counterweight, and most entries do not carry one.
Reversibility and readable outcome
Potential says an idea would matter. It says nothing about whether a person could learn anything by trying it, which is a separate question with a mostly separate answer. Some entries carry a second block covering the two properties that decide it.
Can you undo it? Three values — reverses on stopping, slow to reverse, potentially permanent — each with the specific mechanism written out. This is a claim about the compound's failure mode, not about dose or risk tolerance. Suppression of endogenous hormone production is permanent-shaped in a way that a long half-life is not, and the note says which one applies.
Can you tell if it worked? What a person could actually observe, and over what timeframe. Where the honest answer is nothing, the block says nothing you can measure rather than leaving the row out. That is the most useful line on the page for a large share of frontier compounds: an experiment with no readout cannot inform you however it turns out, and the absence of an endpoint is easy to mistake for an oversight rather than the finding it is. Each entry also names the specific thing most likely to fool someone reading their own result.
These fields describe the experiment. They are deliberately not combined into a score, a ranking, or any flag meaning "safe to try" — a compound can be fully reversible and cleanly measurable and still be a bad idea, and no arrangement of these two facts licenses that conclusion. Nothing here is medical advice.
What the score isn't
It's a running synthesis, not a prescription. Four big caveats:
- It is not universal. A result in people with a diagnosed condition does not automatically apply to healthy people, and a biomarker change is not automatically a longer or better life.
- Individual applicability is unresolved. Population-level findings do not determine an individual's likely outcome. Health history, genetics, other medications, formulation, and monitoring can change both benefit and harm.
- Dose and context alter the evidence. Effects and harms can differ by dose, route, formulation, duration, and indication. The dose field records studied, labeled, clinical-practice, or reported-use contexts; it is not a protocol.
- Evidence and corrections move. Grades update as trials land, sources are re-checked, and errors are found. A “last reviewed” date records an editorial pass, not a guarantee that every citation was independently replicated. Material changes belong in the public archive.
Who does the scoring
I do. I'm not a clinician. I read the literature and call scores based on what I read. The project does not currently claim independent medical review. Sources are linked so readers can inspect the basis of an entry, challenge its scope, and report mistakes. Substantive corrections are dated publicly.
Full context: who I am · disclaimer.