Methodology

How the score is built.

Every intervention gets three 0–10 axes—Evidence, Benefit, and Safety— blended into a weighted composite, then mapped to a 14-grade letter scale. The same arithmetic is used across categories, but each score only applies to the population, use, outcomes, dose, and duration described in that entry.

The unit being scored

Evidence does not belong to a molecule in the abstract. It belongs to a claim: an intervention used by a defined population, for a defined purpose, compared with something else, over a stated period. Vitamin D for confirmed deficiency is not the same claim as vitamin D for preventing disease in an already-sufficient adult. The current scorer uses the first clause as the headline score, so that clause must name the exact claim and population it represents. Other contexts remain visible in the entry; mixed-use rows are being migrated toward separate context-specific claim blocks.

The formula

Composite = 0.45·Ev + 0.15·Bn + 0.40·Sf
(each axis 0–10 → composite 0–10)

Evidence carries the heaviest weight because the brand promise is evidence-first. Safety is next because a large effect with bad tolerability or serious downside should be pulled down visibly. Benefit still matters, but it cannot dominate the grade on its own. A numeric composite is descriptive, not a recommendation: an intervention that is effective and dangerous can still produce a mid-range number. Entries with a high-risk evidence profile are labeled so readers do not mistake the composite for a safety score; the Benefit and Safety axes remain separate and equally visible. Cost and legal access aren't scored at all: they vary too much by country, insurer, and supplier to belong in a universal grade. Each compound page still notes its legal status as context — it just doesn't move the letter grade.

Evidence (45%)

How much is actually known about this in humans?

Study design alone does not settle certainty. Risk of bias, consistency, directness, precision, registration, attrition, outcome choice, and publication bias all matter. Conference abstracts, preprints, registry results, and company toplines are labeled provisional; they should not silently receive the same weight as a peer-reviewed trial or regulator-reviewed evidence package.

Benefit (15%)

If the claim is correct, how big is the effect?

Safety (40%)

Inverted from risk: higher score = safer.

What about cost and legality?

Neither feeds the score. How hard a compound is to obtain — prescription vs. over-the-counter, cash price, schedule status — depends on where you live, your insurer, and your supplier, so it can't sit inside a universal grade without distorting it. Each compound page records its legal status as a note for context; it just doesn't change the letter grade.

From composite to letter grade

The 0–10 composite gets mapped to a 14-grade letter scale. Most grades cover a 0.6-point band; the top four (A− through A+) are tightened to 0.3 points because the best-evidenced interventions cluster near the top of the scale. S covers everything from 8.5 up; F catches everything below 2.2.

S ≥ 8.5 Highest composite range: robust evidence, meaningful effect, and a high Safety score in the stated context.
A+ 8.2 – 8.5 Very high composite: converging evidence with strong Effect and Safety axes.
A 7.9 – 8.2 High composite: strong evidence with meaningful effect and a comparatively high Safety score.
A- 7.6 – 7.9 Strong with one weaker axis (modest effect size or a thinner evidence base).
B+ 7.0 – 7.6 Upper-middle composite: most axes are high, with one lower axis.
B 6.4 – 7.0 Middle-high composite: moderate-to-strong evidence with a modest effect or lower Safety axis.
B- 5.8 – 6.4 Middle composite with an axis profile that depends substantially on context.
C+ 5.2 – 5.8 Limited. Mixed evidence, small effect, or noticeable risk.
C 4.6 – 5.2 Limited. One or two real weaknesses.
C- 4.0 – 4.6 Lower-middle composite: one or more axes are low, narrow, or uncertain.
D+ 3.4 – 4.0 Low composite: thin data, small or unclear effects, or a low Safety score.
D 2.8 – 3.4 Weak. Multiple axes underperform.
D- 2.2 – 2.8 Weak lower edge. Evidence, effect, or safety is especially low.
F < 2.2 Lowest composite range: one or more axes are absent or near the bottom of the scale.

Toxins use a separate scale

The toxin model (heavy metals, pollutants, and lifestyle exposures) is designed to characterize evidence of harm and exposure rather than an intervention-benefit axis. It reports exposure concern: harm magnitude, evidence of that harm, and exposure prevalence. A higher score describes a stronger and/or more widespread harm signal; it is not an instruction for an individual reader.

Exposure concern = 0.40·Magnitude + 0.30·Evidence + 0.30·Prevalence

CRITICAL ≥ 8.5 Population-level harm, strong evidence, and broad exposure.
HIGH ≥ 6.5 Significant harm and/or broad exposure.
MODERATE ≥ 4.0 Evidence-backed harm in some settings.
LOW ≥ 2.0 Limited or context-dependent harm.
MINIMAL ≥ 0.0 Weak or speculative harm signal.

Frontier potential — a second, separate number

The composite answers one question: how well established is this? On most entries that is the only question worth asking. On the frontier it is not.

An example from our own database. Epithalon, a pineal peptide resting largely on one research tradition, graded C+. Rapamycin, the most reproducibly life-extending compound in animal biology, graded C. The formula produced that ordering honestly: Safety is 40% of the composite, and Epithalon scores high on Safety because almost nobody has studied it. Rapamycin scores lower because it has been studied enough to document real immunosuppression. A high Safety score earned by absence of scrutiny outweighed one earned under it.

The composite was not wrong about what is established — rapamycin has no proven human longevity benefit. It was simply the only number on the page, so it got read as a measure of how interesting something is. Rather than reweight the formula and silently re-grade every entry, some entries now carry a second number alongside the grade.

What it is

A 0–10 judgment answering a different question: if this worked in humans, how much would it matter? Rapamycin scores 9 on that axis while grading C on evidence, and both numbers are true at once.

Three rules that keep it honest

What it is not

It is editorial judgment, labelled as such wherever it appears, and it is the least defensible number on this site — unlike Evidence and Safety, it does not come from documents. It is a statement about an idea, not a prediction that any reader will benefit, not a recommendation, and not a substitute for the evidence grade sitting next to it. Coverage is deliberately partial: it appears only where we can write a real basis and a real counterweight, and most entries do not carry one.

Reversibility and readable outcome

Potential says an idea would matter. It says nothing about whether a person could learn anything by trying it, which is a separate question with a mostly separate answer. Some entries carry a second block covering the two properties that decide it.

Can you undo it? Three values — reverses on stopping, slow to reverse, potentially permanent — each with the specific mechanism written out. This is a claim about the compound's failure mode, not about dose or risk tolerance. Suppression of endogenous hormone production is permanent-shaped in a way that a long half-life is not, and the note says which one applies.

Can you tell if it worked? What a person could actually observe, and over what timeframe. Where the honest answer is nothing, the block says nothing you can measure rather than leaving the row out. That is the most useful line on the page for a large share of frontier compounds: an experiment with no readout cannot inform you however it turns out, and the absence of an endpoint is easy to mistake for an oversight rather than the finding it is. Each entry also names the specific thing most likely to fool someone reading their own result.

These fields describe the experiment. They are deliberately not combined into a score, a ranking, or any flag meaning "safe to try" — a compound can be fully reversible and cleanly measurable and still be a bad idea, and no arrangement of these two facts licenses that conclusion. Nothing here is medical advice.

What the score isn't

It's a running synthesis, not a prescription. Four big caveats:

Who does the scoring

I do. I'm not a clinician. I read the literature and call scores based on what I read. The project does not currently claim independent medical review. Sources are linked so readers can inspect the basis of an entry, challenge its scope, and report mistakes. Substantive corrections are dated publicly.

Full context: who I am · disclaimer.