< JMS / ARGUMENTS

[ THE ULTIMATE GUIDE ]

The AI
Question

Every serious argument, for and against building more capable AI, side by side — and every one of them opens, with a steelman underneath from whoever takes the other side seriously. Underneath that: the honest spread of expert opinion. Start here, though: why we're building this at all.

[ 01 — WHY WE BUILD ]

The reason isn't the chatbot.
It's what's underneath it.

"AI" now means a search box that writes emails, which is a genuinely small part of why anyone serious is pushing this forward. Five reasons the answer to "should we keep building more capable AI?" starts as yes — and then the rest of this page is the argument over how carefully, and on what terms.

01 Biology stopped being the bottleneck // a fifty-year problem got solved by a model, and medicine noticed
2024 Nobel Prize in Chemistry for AlphaFold 190+ countries using it, 2M+ researchers 87.2% sensitivity, first autonomous AI diagnostic device

Protein structure prediction ate fifty years of biology PhDs before a model cracked it. In 2024, the Nobel Prize in Chemistry went to Demis Hassabis and John Jumper "for protein structure prediction" and David Baker "for computational protein design" — the AlphaFold work. By then it had already been used by more than two million researchers in over 190 countries, on problems from neglected tropical diseases to enzyme design that would never have justified a standalone research grant.

It's not an isolated case. In April 2018 the FDA cleared IDx-DR, the first autonomous AI system permitted to diagnose a medical condition — diabetic retinopathy — without a specialist reviewing the result, at 87.2% sensitivity and 90.7% specificity in its pivotal trial. That's a screening exam a primary care office can now run on the spot instead of a referral and a wait. Drug-discovery pipelines built on the same techniques are now advancing AI-designed candidates into human trials. None of that required a general chatbot. All of it required the underlying capability to keep improving.

02 Expertise gets cheap for the first time // the same capability that drafts a contract also explains one for free

For most of history, good analysis, careful writing, and patient explanation were rationed by what you could afford. A small business without a lawyer on retainer, a student without a tutor, a clinic without a specialist down the hall — all of them previously did without, or did worse.

A model that can explain a lease, draft a first pass at a business plan, translate a school notice, or walk a kid through the calculus problem their district can't staff a teacher for doesn't replace any of those professionals. It replaces doing without. That's a genuinely different distribution of who gets to make a well-informed decision, and it scales in a way that hiring more lawyers and teachers never has.

03 The productivity story is already in the numbers // not a forecast anymore — it's showing up in the accounts
~7% projected global GDP lift, Goldman Sachs ~$7T over a decade ~25% of labor tasks automatable, advanced economies

Goldman Sachs estimates generative AI could raise global GDP by roughly 7%, on the order of $7 trillion, over a ten-year horizon, with productivity growth as the engine. The same research puts perhaps a quarter of labor tasks in advanced economies within reach of some degree of automation over that period.

That's a genuinely different order of economic event than most technologies produce, and it arrives at a moment when productivity growth has been sluggish for two decades and workforces in several major economies are shrinking rather than growing. A tool that lets existing workers produce more, not just cheaper, is one of the few realistic paths to sustained higher living standards without either population growth or a painful reallocation of who has to work harder.

04 The same model that writes your email tracks the hurricane // general-purpose means the benefit isn't limited to chat

The consumer product is the visible ten percent. The same underlying techniques are running probabilistic weather models that give forecasters more lead time on a storm track, screening candidate battery and superconductor chemistries without a chemist at a bench, modeling plasma confinement for fusion reactors, and planning the transmission upgrades a grid has been deferring for decades.

None of that is speculative. It's existing technique that got faster and cheaper as the underlying models got more capable — and it keeps getting more capable because the same research investment that makes the chatbot better also makes the hurricane model better. The two aren't actually separable, however separately they get talked about.

05 The alternative to building is not "nobody builds it" // and safety needs the capability to see clearly, not less of it

The most common mistake in this debate is treating "stop building" as the neutral, safe default. It isn't one of the options on the table. Frontier AI research is currently happening at a small number of labs in a small number of countries, and every major government developing it has said explicitly that ceding the frontier to a less careful competitor is not a plan they'll accept. If the most safety-conscious labs slow down unilaterally, the field does not pause — it moves to whoever slows down the least.

There's a second reason capability and caution aren't actually opposites: the tools used to evaluate, red-team, and interpret an AI system get better as the underlying models get better. Understanding what a system can do, and catching what it shouldn't, is itself a research problem that benefits from more capable AI working on it — which is exactly why the rest of this page isn't a sales pitch for speed. It's every argument, both directions, with the steelman attached.

[ 02 — THE DIVIDE ]

Now the honest part: how worried should you actually be?

Both sides of this argument have an incentive to sound more certain than they are. Here is the closest thing to a real measurement — a 2023 survey of 2,778 researchers who had published at top AI venues, run by AI Impacts. The question: what's the probability advanced AI leads to an outcome as bad as human extinction?

AI_RESEARCHER_SURVEY // AI Impacts, Fall 2023 · 2,778 published AI researchers

// The median answer was 5%. Depending on exactly how the question was phrased, the mean ran roughly 9% to 16% — a wide gap from the median that, by itself, tells you the distribution has a long, worried tail.

68%

of the same researchers said good outcomes are more likely than bad ones overall.

48%

of those optimists still gave at least a 5% chance to an extremely bad outcome.

59%

of the net pessimists gave at least 5% to an extremely good outcome right back.

Read that carefully, because it's the actual finding, not the headline version. There are almost no confident people in either direction. High hopes and real dread mostly show up in the same person — which is a more honest starting point than picking a side and looking for researchers who agree with you.

"We've never had to deal with things more intelligent than ourselves before. And how many examples do you know of a more intelligent thing being controlled by a less intelligent thing?"

— Geoffrey Hinton, 2024 Nobel laureate in Physics for foundational neural network research, who has put the chance of AI causing human extinction at roughly 10–20% within 30 years (BBC Radio 4, Dec 2024)

[ 03 — THE LEDGER ]

Every argument. Both sides. Open them all.

Forty-two arguments, matched left and right. Every card opens into the actual claim plus the steelman — the strongest response the other side has to that specific point, written by someone who takes it seriously. Filter by topic, search the text, and mark which arguments actually move you. The scale keeps score.

A balance scale weighing the arguments you have marked as mattering

Mark arguments as MATTERS or DECISIVE as you read and the scale will tip. Nothing weighed yet — the beam is level.

▲ THE CASE FOR

▼ THE CASE AGAINST

[ 04 — THE HONEST VERDICT ]

The question was never yes or no.

Read all forty-two arguments and a pattern shows up: almost nobody serious on either side is actually arguing for zero AI research or for zero caution. The real disagreement is about sequencing and enforceability — whether the safety work happens before or after the capability ships, and whether anyone can actually check. Four conditions decide which future you get out of the same underlying technology.

TERM 01

Fund the eval like it's the launch

Bio, cyber, and autonomy uplift evaluations happen before a capability ships, not as a postmortem after something goes wrong. Several frontier labs already run this as a gating process rather than a courtesy check — that should be the floor, not the exception.

TERM 02

Verification beats good faith

A safety commitment nobody can check is a press release. Compute reporting, external red-teaming, and safety institutes with actual audit access turn "we're being careful" into something a skeptic can confirm — the same shift ratepayer protection needed before it meant anything.

TERM 03

Spread the upside on purpose

Diffusion of frontier capability to hospitals, schools, small businesses, and open research doesn't happen automatically just because the technology exists — it happens because someone decided concentration was a problem worth designing against, the same way it was for cloud computing and mobile chips before it.

TERM 04

Write the off-ramp before you need it

Define the red line and the rollback plan before deployment — what result triggers a pause, a restriction, or a withdrawal, and who has the authority to pull it. A capability that can't credibly be paused was never actually being tested, just announced.

Get this right and it's the best tool humanity has ever pointed at disease, poverty, and its own ignorance. Get it wrong and the mistake doesn't stay contained to whoever made it.

Both of those outcomes come from the same underlying technology. That's the whole argument. The capability isn't the variable — whether the safety work is funded, checked, and enforceable at the same pace as the capability is.

So be relentlessly pro-progress and relentlessly pro-safety at the same time. They have never actually been opposites, and treating them as opposites is how you end up with neither the benefits nor the guardrails. The honest answer to "should we build this" was always "yes — and the terms are the whole ballgame."

[ SOURCES + METHODOLOGY ]

// EXPERT OPINION & RISK ESTIMATES

  • AI Impacts — "Thousands of AI Authors on the Future of AI" (2023 Expert Survey on Progress in AI), 2,778 respondents who had published at top AI venues: median 5% / mean roughly 9–16% (varying by question phrasing) probability of an extremely bad outcome, e.g. human extinction; 38% gave ≥10%, 10% gave ≥25%, 1% gave ≥75%; 68% net optimistic on outcomes overall, with 48% of those optimists still giving ≥5% to an extremely bad outcome and 59% of net pessimists giving ≥5% to an extremely good one.
  • Geoffrey Hinton — BBC Radio 4 interview, December 2024: 10–20% probability of AI causing human extinction within roughly 30 years. Hinton shared the 2024 Nobel Prize in Physics with John Hopfield "for foundational discoveries and inventions that enable machine learning with artificial neural networks."

// HEALTH & SCIENCE

  • The Nobel Prize in Chemistry 2024 — awarded to David Baker "for computational protein design," and jointly to Demis Hassabis and John Jumper "for protein structure prediction" (AlphaFold2); used by more than two million researchers across 190+ countries by October 2024.
  • U.S. Food and Drug Administration — April 2018 marketing authorization for IDx-DR, the first autonomous AI diagnostic device, for detecting diabetic retinopathy without specialist over-read; pivotal trial (900 patients, 10 US primary-care sites) reported 87.2% sensitivity, 90.7% specificity, 96.1% imageability.

// ECONOMY

  • Goldman Sachs — "Generative AI Could Raise Global GDP by 7 Percent" (2023): roughly 7% (~$7 trillion) potential lift to global GDP over a 10-year horizon; automation exposure estimated at roughly 25% of labor tasks in advanced economies, 10–20% in emerging economies.

// ENVIRONMENT

  • Li, Yang, Islam & Ren, "Making AI Less 'Thirsty'" (2023) — training GPT-3 in Microsoft's US data centers consumed an estimated 700,000 liters of onsite water, ~5.4 million liters including offsite electricity generation; roughly a 500ml bottle of water per 10–50 medium-length inference exchanges, depending on time and location.

// FISCAL & MILITARY

  • 2025 federal tax law (H.R.1, the reconciliation act commonly called the "One Big Beautiful Bill Act") — 100% permanent bonus depreciation restored for property with a useful life of 15 years or less, applicable to data center and AI infrastructure equipment.
  • US Department of Defense / Defense Innovation Unit — Replicator initiative announced August 28, 2023, targeting thousands of attritable autonomous systems by August 2025. Congressional Research Service — reporting only hundreds, not thousands, of systems fielded by that target date.

On honesty: the steelman under each argument on this page is written by the site, in the voice of someone who takes that side seriously — it is not a verbatim quote from a named person unless directly cited as one. Where a number has a real, checkable source, it's cited above; where an argument is philosophical or predictive rather than measured, it's presented as an argument, not a statistic. Weights you set in the ledger are stored in your own browser and go nowhere.

// jeremymsparks.com — last reviewed August 2026