Proposed Reforms
What we don't actually know

The studies that would settle this

Most pages on this site end with a caveat: a range that's too wide, a follow-up period that's too short, a sample that's too small or too self-selected to lean on hard. Those caveats usually point at a specific, missing study — not a vague call for "more research." This page collects them in one place: what we don't know, what a study that actually answered it would need to look like, and who's positioned to run it.

What this page is not
  • Not a claim that the evidence is too weak to act on. Elsewhere on this site we argue the honest reading of current evidence already favors waiting for consent.
  • Not a wishlist of studies that would flatter one side. Several of these if done well could easily cut against the position this site takes.
What it is

A companion to the Proposed reforms page. The reforms page asks for procedural changes achievable now; this page asks a different question: where are specific points where the data genuinely runs out.

1

A standard definition of the procedure itself — and metrics for when it's gone wrong

This sounds like it should already exist. It doesn't, not in any usable form. "Circumcision" is treated in the literature and in consent paperwork as a single, uniform procedure, but what actually happens on the table varies enormously: how much outer skin is removed, whether the frenulum is preserved, reduced, or taken entirely, how much of the inner mucosa goes with it, how tight the resulting shaft-skin is left, how much is judged "too much" versus "not enough." None of that is standardized, and (this is the actual gap) there isn't an agreed set of outcome metrics that would let anyone say, after the fact, that a given result was appropriate versus botched.

That absence matters for the debate itself, not just for individual cases. Complication studies mostly count things that are unambiguous — bleeding, infection, need for revision surgery — because those are the outcomes existing coding systems can capture. A result that falls short of a surgical emergency but still removed more than a functional amount, or left a degree of scarring or skin tightness that impairs later sexual function, has nowhere to be recorded. It isn't in the complication rate, because clinically nothing "went wrong." Whether an amount of tissue loss inside that normal range is itself a harm worth counting is exactly the question this site can't answer with a number — and neither, currently, can anyone else.

It also quietly undermines proposal 3, below. A sensation or satisfaction study that recruits "circumcised men" as one undifferentiated group, without recording how much was actually removed, can end up sampling disproportionately from one end of that range — say, men with a preserved frenulum and more residual inner mucosa — and then reporting a null result for the whole category, including men with the frenulum removed and little to no inner skin left. Those are not the same procedure wearing the same name. Without a shared way to record what was actually done, no sensation study can say whether a finding is about circumcision in general or just about the specific, unrecorded version of it that happened to make up the sample.

What it would take

A validated, standardized measurement and grading system for circumcision outcomes — analogous to grading scales already used elsewhere in surgery — that records how much outer skin and inner mucosa were removed relative to a defined baseline, frenulum status, residual shaft-skin mobility, and cosmetic outcome, collected prospectively across providers and techniques (clamp, Plastibell, freehand). Only once that exists does "botched" stop being a subjective judgment call and become a measurable deviation from a stated range — and only then can anyone meaningfully ask how often outcomes fall outside it.

Who's positioned to run it: pediatric urologists and surgical outcomes researchers, ideally the same professional societies that already publish technique guidelines — the gap is a measurement standard, not a lack of surgeons willing to be studied.

2

A lifetime risk model that includes therapeutic circumcision

Every complication comparison in the literature treats "circumcised" and "intact" as two fixed, separate populations with their own flat complication rates. That understates the real comparison, because a meaningful share of intact complications (severe phimosis, recurrent balanitis, paraphimosis) are ultimately resolved by circumcision, performed later, on a less favourable timeline, often under general anaesthetic instead of a local block.

A newborn circumcised at birth pays that surgical-risk cost once, upfront, for the whole population. An intact boy only pays it if he later develops one of the indications that actually leads to circumcision — and when he does, the procedure itself may carry a higher complication rate than the neonatal version being compared against. Neither side of that trade is currently modelled anywhere in the literature this site could find.

What it would take

Not a new trial — a modelling study, built from data that partly already exists: the proportion of intact boys who ever receive a therapeutic circumcision (by age and indication), and complication rates for circumcision performed at each later age band versus the neonatal baseline. Combined, those numbers would produce an actual lifetime expected-complication comparison, instead of two disconnected snapshot rates.

Who's positioned to run it: health economists or pediatric urology researchers with access to a large insurance-claims or national-registry dataset — the same kind of database already used for large-scale complication studies elsewhere in the literature.

3

A sexual-function study that isn't small, self-reported, or confounded

Put plainly: nobody has won this outright. The measurement studies (fine-touch thresholds, quantitative sensory testing) suggest the removed tissue is genuinely the most sensitive part of the organ. The studies that try to translate that into lived sexual experience and satisfaction are, on both sides of the argument, small, self-reported, cross-sectional, or drawn from culturally confounded populations. The same bar belongs on the sexual-function literature around female genital cutting, which draws identical criticisms (small, self-reported, confounded), so that neither side of the broader genital-cutting debate gets held to a looser evidentiary standard than the other.

What it would take

A study that pairs objective quantitative sensory testing with validated, standardised sexual-satisfaction instruments, done prospectively rather than as a retrospective survey, in a population large and diverse enough to separate the effect of circumcision from the effect of the culture that made the decision — and one that records how much tissue was actually removed for each participant (frenulum status, residual inner mucosa) rather than treating "circumcised" as a single category. See proposal 1: without that data, a study built on a sample skewed toward one end of that range can't tell you what it thinks it's telling you about the other end.

Who's positioned to run it: sexual medicine researchers with access to a natural-experiment population — adult-circumcision cohorts, or countries with mixed intact/circumcised populations not sorted by religion, are the likeliest source of a design that isn't hopelessly confounded.

4

A real, prospective adverse-event registry that takes reports from adults

This one is already argued in full on the Proposed reforms page — it's proposal 3 there. The short version: the best complication numbers we have were reconstructed decades after the fact, from records nobody built for this purpose, because structurally nobody is required to track circumcision complications the way vaccines or blood transfusions already are. "If complications were common, we'd already know" only holds if someone is looking.

What it would take

A mandatory or opt-in complication-reporting registry, prospective rather than reconstructed, with a self-report pathway for adults — because some of the outcomes this debate is actually about (sensation, sexual function, psychological effects) surface at puberty or later, long after a newborn-period clinician has stopped watching and feedback is difficult to give.

Who's positioned to run it: professional societies, health regulators, or hospital systems — see Proposed reforms for the full case.

5

A longitudinal, cross-cultural study of psychological outcomes

The strongest evidence the Trauma page has for lasting psychological effects (Miani et al., 2020) is a real, peer-reviewed finding, and the authors are honest about its limits: cross-sectional, self-reported, drawn entirely from a US population where circumcision status correlates with religion and region in ways that could themselves explain personality differences. It cannot establish that circumcision came first, in any causal sense.

What it would take

A prospective cohort followed from infancy into adulthood on standardised psychological measures, ideally spanning a population — like a country with a genuine intact/circumcised mix not sorted by religious practice — where the confound the current study can't rule out simply doesn't apply.

Who's positioned to run it: developmental psychologists running existing birth-cohort studies, who could add circumcision status and the relevant instruments to a cohort they're already tracking for other reasons.

A site that only ever tells you what the evidence shows, and never what it can't yet show, is quietly overselling its own certainty.

If any of these five proposals moved forward tomorrow, some of them might strengthen the case this site makes. Some might weaken it. That's what makes them worth wanting — not because we're confident how they'd come out, but because right now nobody actually knows. Considering how common the procedure is and how permanent it is, we owe it to boys to actually know these things.