AI companies are racing to build AIs that are smarter than humans in every way. In AI 2027, we predicted that this would result in either extinction or irreversible concentration of power.1
Plan A is our positive vision for what should happen instead.
In this scenario, humanity delays the development of superintelligence until 2040, makes all AI research public, allows dozens of companies globally to catch up to the frontier, and intentionally enters a regime of mutually assured compute destruction.
Plan A is our positive vision for how humanity can avoid AI-driven existential catastrophe and reach a flourishing future. It’s informed by conversations with experts at major U.S. frontier AI companies, direct experience at OpenAI, tabletop exercises, and discussions with policymakers, national security experts, and AI policy leaders. We recommend an international deal to avoid a dangerous race to superintelligence. The deal involves total research transparency for AI R&D, which allows the nations of the world to understand what’s happening and enforce guardrails. The result is multiple companies across multiple countries scaling slowly and safely together towards superintelligence, instead of racing each other in secrecy.
Plan A is primarily a recommendation, not a prediction. This scenario is not our best guess as to what the future will actually look like. Instead, it’s a vehicle for communicating and stress-testing our policy recommendations. While the implementation of Plan A is a recommendation and not what we actually expect to happen, the subsequent effects depicted are predictions.2
In this AI 2040 scenario, Plan A is implemented successfully, albeit imperfectly and only in the nick of time.
We contrast Plan A with 4 alternative plans (B, C, D, and S), which correspond to the main ways the US could respond (or not) to the challenges of superintelligence.3
AI companies will probably succeed at their stated goal of building smarter-than-human AI systems within the next 1 to 10 years.
The industry has convinced itself that controlling superintelligent AI can be figured out on the fly, and thus has no remotely adequate plan. We think this situation is terrible and could easily get us all killed.4 We do not expect whoever “wins the race” to have much of a lead, and we do not expect them to unilaterally slow down to reduce existential risk.5 If this race continues,6 we do not expect humans to maintain effective control as their AIs become superintelligent.7
Moreover, even if the AI companies somehow align their AIs, the result will be an unprecedented concentration of power—that is, the result will be a situation where a tiny group of people, or possibly just a single individual, is effectively in control of the world’s only army of superintelligences for some months, and will be presented by said superintelligences with various options for how to proceed, some of which will de facto amount to taking over the world.8
As best as we can guess, the CEOs of OpenAI, Anthropic, xAI, and Google DeepMind understand this and are proceeding anyway, perhaps because they think they are the lesser evil and will use their immense power responsibly, unlike Xi Jinping or rival CEOs.9
While we agree that it is generally correct to choose the lesser evil, we don’t think we should advocate for a strategy that has such a scarily high chance of leading to human extinction or global dictatorship. Instead, we wish to advocate for something that is actually good. If enough people do likewise, it can happen.
So, we wrote a scenario outlining that possible world.
“Plans are worthless, but planning is everything.” - Dwight D. Eisenhower
We think most AI policy proposals fall apart under scenario scrutiny—that is, if you try to write down a detailed and plausible scenario in which that proposal succeeds, you will find it difficult to do so, and you will realize the plan is less likely to work than it seemed, or has more unpleasant side-effects than its proponents acknowledged.
Perhaps that’s why scenario scrutiny is so rare in AI policy. Everyone wants to say that their own favorite policies will have great consequences and that the policies of their rivals will have terrible consequences. Applying scenario scrutiny to their own favorite policies might surface uncomfortable issues with them; meanwhile, applying scenario scrutiny to their rival’s policies is a lot of work for little rhetorical gain.10
We think the discourse would be improved if more AI policy proposals were subjected to scenario scrutiny. So we’re starting with our own, even though this opens us up to criticism. We hope critics will judge us against the existing state-of-the-art for plans to navigate the AI transition (if they can find any) and not against some hazy but pleasant fantasy where no one has to make any hard choices yet everything will probably be fine.
What of the immense difficulty of predicting the effect our policy would have in a world approaching superhuman AIs? This is like trying to predict how to best fight World War 3, except that it’s an even larger departure from past case-studies. Yet it is still valuable to attempt, just as it is valuable for the U.S. military to game out Taiwan scenarios in excruciating detail. There are other precedents as well: intelligence agencies, climate bodies, and pandemic-preparedness offices all rely on various kinds of scenario planning.
Plan A is our ambitious proposal for what to do, and we’d like to see something like it implemented soon because we are uncertain about how much time remains.11 But for purposes of writing a concrete scenario, we need a concrete timeline.
The timeline of this scenario is:
In 2029, the US and China agree to avoid a reckless race to superintelligence.
In 2030, we would have fully automated AI R&D, leading to superintelligence by the end of the year. Thanks to the deal, we avoid this.
Between 2030 and 2035, we scale within the human range, to AIs that are roughly as capable as top human experts.
In 2035, we pause at top-human-expert level AI in order to maintain human control.
In 2040, we unpause and scale to superintelligence.12 (Hence the title: AI 2040)
In our previous scenario, AI 2027, AI fully automated the process of building smarter AIs in 2027, leading to an intelligence explosion and superintelligence within the year. The two differences in this scenario are (1) the default timeline is now 2030, and (2) thanks to governance actions, generally-superhuman AIs first appear in 2040.
We changed the default timeline because we want our portfolio of scenarios to reflect our uncertainty about AI timelines. AI 2027's titular year was chosen because, at the time we started writing, Daniel thought there was roughly a 50% chance that things would go that fast or faster.13 At the time we started writing Plan A, 2030 was the corresponding year for Thomas. Daniel currently thinks things will probably go somewhat faster than depicted in this scenario; you can read more about our team’s views on timelines here and here.
We changed the governance actions because this scenario is primarily a recommendation, not a prediction. Conducting a full-speed intelligence explosion is wildly reckless and concentrates power to an extreme degree.
America has two workforces now. The first is people, 165 million of them. The second is AI agents: millions of copies spun up and shut down every hour, working around the clock at superhuman speeds.
Most of their work is slop. But enough of it is good that people are paying ten billion dollars a month for AIs that can, in theory at least, do anything on a computer that an employee can.
There is one job the AI companies want to automate more than any other—their own. They haven’t succeeded yet; no recursive self-improvement so far.14 But they seem to be getting closer, and they’re pulling up the ladder behind them: the strongest coding AIs refuse to help competitors with AI R&D.15 Even as the most bullish employees admit that things are taking a bit longer than planned, the skeptics notice that their usual dismissals are starting to ring hollow. Why exactly will AI never be able to do my job? What’s the barrier again?
Congress is starting to pay more attention. They’ve long been hearing about AI: datacenters using too much water,16 chatbots encouraging suicide, Mythos hacking NSA systems—and of course, tech industry lobbyists warning that any whiff of regulation will make America immediately lose the race with China and spend the rest of history as a CCP tributary state.17
Now they step back and ask: Where are we going with this? What does the world look like five, ten, or fifteen years from now? Will there still be jobs? What if there aren’t?
One question weighs especially heavily on their minds: Who will control all these AIs?
Congress settles on an important part of the answer: Probably not us.18
They hold a series of tense hearings on AI. They read the 2016 OpenAI emails discussing how OpenAI was founded in order to prevent Demis Hassabis from becoming dictator.19 But who is preventing Sam or Elon from becoming dictator? Congress is unsatisfied with existing responses.
The result of this wakeup is the AI Transparency Act of 2027, an omnibus bill that does many things, some good and some bad, but doesn’t fundamentally change the situation.20
Our main recommendation is to begin negotiating something like Plan A as soon as possible. But in this scenario, we depict Plan A happening imperfectly and only in the nick of time. So here is a list of less ambitious ideas that still help.
The 2028 election cycle is heated, as usual. AI is the biggest topic. The datacenters now under construction cost twice as much as the entire US military budget.23
Most white-collar professions are seeing disruption like software engineering saw in 2026; such jobs now heavily involve managing AI agents. AI companies have industrialized the training process: Executives say “let’s move into [profession] this year” and then the company interviews professionals, buys data, creates training environments, etc. until their AIs get traction. Then the AIs rapidly improve as they are used more widely in the field and accumulate more real-world data.
Other countries are starting to get scared and angry. It seems like a handful of US and Chinese companies are on track to automate all the white-collar jobs. Power is concentrating in the US, and in particular in the President plus a handful of tech CEOs.
AI experts warn that the intelligence explosion is near. By speeding up AI research, the AIs will become even more competent, speeding up research even faster, making them even more competent, and so on. There are complicated dynamics about bottlenecks and hardware limits governing how fast this process goes and where it ends, but it seems like it might go very fast and end somewhere very far away.
On the default path, the next presidential term will see AIs that are far beyond human level, created entirely by AIs, themselves created entirely by other AIs, without any human in the loop since several generations back. Will those AIs be obedient, aligned, etc.? Why? Who will control them if so? How exactly is all of this supposed to end well?
Having put humanity on this path, the AI companies find it acceptable. But most people don’t. Forget thinking about his legacy—the President is starting to think about what’ll happen to him after he leaves office and the world gets transformed.24 Both presidential candidates keep getting asked what they’ll do about AI, and try out increasingly dramatic ideas on the campaign trail. The discourse bounces back and forth across all of the options displayed below, and more.
Eventually the President and his protégé converge on one plan; the opposition candidate converges on another. Then it’s Election Day.
“Trust, but verify” – Ronald Reagan
The President announces that the US will pursue international cooperation to avoid an imminent intelligence explosion.
“This mad race toward superintelligence must end. For too long, we have been pursuing lesser-evils and least-bad solutions. We need a Plan A. We can still proceed with AI development, but we must do it more cautiously, more transparently, and involving many more countries and companies.”
To the surprise of many in DC, China proves receptive. They had been debating the same issues—social destabilization, job loss, rogue superintelligence - on their side of the Pacific. They had been looking forward to “The Chinese Century” and thought that AI might disrupt their plans. And they had one extra reason to come to the table: the US continues to have more compute, more datacenters, and better models. They were worried about the things the US might do to them if it got to superintelligence first.48 We aren’t confident, but we expect that a deal along the lines of Plan A (described below) would be incentive-compatible for almost everyone, including China. For more on why we think this, read the rest of 2029 and 2030, which explains both what the deal is and who it benefits, and perhaps also read this supplement. In short, the answer is that anyone concerned about loss of control should think Plan A is an improvement, along with anyone concerned about concentration of power—except for the people in whom the power would concentrate by default. For these reasons, we expect strong opposition to Plan A from the leading AI companies. We expect them to rationalize arguments for why the deal is bad and why instead what’s best for America and humanity is a different strategy that just so happens to allow them to continue accumulating massive amounts of power. China, by contrast, is an example of an actor in whom power would not concentrate by default. In 2029 in this scenario, the US has a significant lead in AI capabilities over China and a significant advantage in compute which will compound the lead. The more powerful AI gets, the scarier it will be to fall behind, as AI 2027 and the later years in this scenario illustrate.Why would China be interested in a deal? Why would anyone?
The US and China don’t trust each other. Fortunately, they don’t have to: Plan A includes provisions for verifying compliance. But setting it up will take time.
So for now, they start with something crude. For the rest of 2029, they put a temporary halt to AI training, because that's relatively easy to verify. In 2030, they’ll have the infrastructure in place to proceed with Plan A.
The stock market is gyrating wildly up and down in response to news and commentary about the momentous actions being taken. Everybody is screaming about how it’s going too far or not far enough. The market will calm down by the end of the year, but the screaming will continue.49
Step 1: Compute Declaration
Neither country wants to halt their own AI training unless they can see that the other side is halting too.
Fortunately, training AIs requires large numbers of AI chips. Most AI chips are in giant datacenters.50 AI datacenters are typically big enough to be visible from space, and power-hungry enough to require conspicuous infrastructure. New AI chips can only be manufactured at a handful of fabrication plants (fabs), located mostly in Taiwan, South Korea, the US, and China.
Left: OpenAI’s Stargate datacenter. Right: An Extreme Ultraviolet Lithography (EUV) machine, arguably the most complicated machine ever created by mankind. EUV machines are essential for producing frontier AI chips, and are produced by a single Dutch company, ASML.
The US and China negotiate with the countries that have a major role in the chip supply chain, and they require each major datacenter owner (and their upstream suppliers, including chip fabs) to publicly declare their major purchases and sales.51 Analysts from numerous countries and institutions may pore over the list, ask questions, and challenge anomalies. The rival powers also send inspectors to each other’s infrastructure, reassuring themselves that the stated numbers are accurate. By the end of the year, each side is confident that the other isn’t hiding more than about 1% of AI compute, and that all potential sources of new compute are monitored and accounted for.52
Step 2: Pause Training
Distinguishing dangerous AIs from safe ones will require more time and understanding, so for now they go with a simple solution: a temporary pause on all new training runs.53
Both sides can still use datacenters for running AIs that already exist (i.e. inference), but they will retrofit each other’s datacenters with devices to verify that they aren’t being used for new training runs.54
The US and China are able to rapidly source enough of these verification devices to retrofit almost every major datacenter.55 The inference-only verification solution we propose involves installing simple network taps and verification servers (that perform partial recomputation) in AI datacenters, but even better solutions may be possible and preferable. Building these devices in advance, in a way that each side trusts sufficiently (at least unilaterally) and that is robust to security vulnerabilities, will be hard and will require early effort. In our scenario, we assume a viable but imperfect version of the devices is ready in advance, and that both sides supplement them early on with enough defense-in-depth measures (e.g., tamper detection) while sprinting to improve their security and robustness. We therefore recommend early investment in verification R&D to improve the options on the table. Ideally, there would be a range of stress-tested solutions ready in advance; and we think it would be extremely cheap relative to total AI investment for such devices to be prepared in advance (i.e., on the order of 0.1% of AI investment, so single digit billions). Even in a situation that is more pessimistic than our scenario, where basically no verification R&D progress has been made since 2026, we still recommend the US and China cooperate to navigate the intelligence explosion slowly and safely. However, in this case, the early period of a deal would have to be more imperfect (i.e., some combination of being harder to verify that the other is complying, or more economically costly, by requiring more AI compute and services to be shut down for a while). The main alternatives to ready-to-go inference-only-devices available are: Rely on intelligence (e.g., spies, satellite monitoring, cyber) to tell if the other side is complying. Implement a temporary software-only version quickly (that relies on pre-existing hardware), that will be less secure but may be sufficient until they upgrade to a better solution. Buy up and install any available off-the-shelf devices as network taps last minute. Without producing more secure ones, this will probably be better than the software-only version but still not very secure. Shut down some fraction of AI compute (e.g., 90%) until better inference-only verification is ready. This would slow R&D substantially but is less painful than it sounds on the inference side: because models are distilled so quickly, roughly 10% of compute suffices to serve the best model from a year earlier. At worst, most users temporarily fall back to year-old AI capabilities. Don’t pause AI capabilities progress, and let it continue for the months it takes to quickly build an inference-only solution and eat the added risk from not slowing down yet.What if they didn’t have inference-only verification ready?
Inspectors from both countries sign off on the installation of the verification devices. Only datacenters that the US and China agree have functioning verification are allowed to provide continuing AI services. Meanwhile, both governments sprint towards more fine-grained verification options that could allow training to restart while keeping both sides confident the other isn’t racing ahead.
Step 3: Get Worldwide Buy-in
Though other countries were consulted from the start, especially those owning parts of the AI supply chain, Plan A began as a bilateral deal between the US and China. Now that it’s taking firmer shape, negotiations go multilateral. Carrots and sticks are waved about.
Many countries outside the US and China are happy about the deal, because they were worried about a future in which a handful of US and Chinese AI companies recursively self-improve, pull farther and farther ahead, and then… Well, what happens next depends on who you ask, but answers range over possibilities such as “take our jobs,” “cement hegemony forever” and “get everyone killed.” So they want the US and China to proceed more slowly and transparently. This will allay their fears and also allow their own AI projects to catch up to the frontier.
So it’s surprisingly easy for the US and China to get buy-in.56 By the end of the year, most of the world has joined what is now called the Consortium.
In our main scenario, neither side attempts to defect from the deal. In this branch, China aggressively cheats. Our forecast is that if Plan A were implemented, they would not acquire enough compute to overtake the consortium. This timeline showcases how that would likely play out.
Tech companies are pushing hard to restart AI training, and lobbying to shape the conditions under which it happens. Large portions of the public are saying that we should never restart. Some countries are too.
After a year of intense discussions, negotiations, and twitter beefs, a compromise takes shape: AI development will continue, but in a more cautious, more transparent, more distributed way.
Plan A is guided by 4 core principles:
Principle 1: Buy Time
The problem with an intelligence explosion is the "explosion" part.
The default trajectory is reckless and destabilizing. Each country and company is racing their competitors for AI dominance without time to think. Most likely, a ton of people will end up disempowered or killed. Maybe godlike AIs run circles around us and the age of humanity is over. Maybe a CEO or President establishes a permanent dictatorship. Maybe World War Three begins in the chaos, as countries without the best AIs fear losing their leverage. Maybe everything is surprisingly smooth and we "only" automate all labor and render everyone unemployable in a radically transformed world. Even aspiring AI-enabled dictators are worried; there's enough frontrunners that none of them have more than 50% odds of winning the power struggle.
No one knows how to tell if AIs are trustworthy, and no one knows how to regulate superintelligent AI. Solving those problems looks like it'll take years. Slowing down the intelligence explosion buys time to do so.
It also helps prevent extreme concentrations of power, because it gives more time for groups that don’t directly control the world’s smartest AIs to wake up to what’s happening before they lose their leverage. Without a major slowdown, a rapid intelligence explosion will concentrate almost all the power into a tiny group of CEOs and politicians. Much has been said about the dangers of going too fast. The dangers of going too slow are less obvious, but still significant. The most pressing is that treaties can’t last forever - for example, the START treaties fell apart after thirty years.73 If Plan A met a similar fate, the race to superintelligence would begin again. Better for it to happen sooner under good conditions, than later during a chaotic race. A second hazard of going too slow is that it might allow defectors enough time to develop destabilizing AI in secret. Auditors have verified the location of 99% of the compute produced before the deal, and see encouraging signs that the remainder is mostly accounting errors and random low-level smuggling. But they cannot rule out that one or both powers have shipped on the order of 1% of world compute to hidden locations where they intend to defect on the deal and race to superintelligence alone. How dangerous would this be? Although it is far harder to train dangerous AI with 1% of world compute than with 10% of world compute (the amount owned by the largest private companies), it’s hard to be highly confident about exactly how hard. (For more analysis of this question, see our Covert AI Projects supplement, or the Covert Project branch above). Although both sides promise that they don’t have covert projects, this is nevertheless another reason to aim for a definitive solution sooner rather than later.74 For more on what an attempt at a complete stop might look like, see our Plan S branch.Why don’t we just completely stop?
Principle 2: Total Research Transparency
So the goal is to develop AI at a reasonable pace: Not too fast, not too slow. Ban the unsafe kinds, allow the safe kinds. But each country still regulates its own AI industry; there isn’t a central planner with authority to regulate AI worldwide. So how do they decide what to allow and disallow, and how do they verify that no one is cutting corners?
The Consortium countries come up with a simple high-level framework: we’ll agree to let each other see all the AI research. Then, if we don’t like something someone is doing, we’ll talk about it and perhaps agree to ban it. Consider the world today. Publishing model weights on the internet is not banned, but it's strongly disincentivised because companies that train billion-dollar models don't want to give them away for free. Open models lag behind and will lag further behind as companies scale up, securitize further, and start automating AI R&D. Now consider the world in which full open source (publishing the weights, the training code, and the data) is mandated for all frontier models. That’s a totally different world, much farther from the status quo than the status quo is from the OS-banned world. Plan A is not that crazy, but it's close: the algorithms are mandated to be open source, and while the weights aren't (in fact the weights are banned from being OS), access to the weights is mandated to be open; members of the public can do evaluations (including fine-tuning!) on frontier models just like employees can (under conditions of total transparency, that is, so they can’t turn around and use the model for actual work). We think that open algorithms plus open access gets most of the benefit of mandating open-weights, without the cost that bad actors can strip off the guardrails and build bioweapons. Why isn’t Plan A more closed? We think the open version of Plan A is more robust and likely to succeed. However, we do think somewhat more closed variants of Plan A are promising: notably Filtered Transparency. Under this proposal, there would still be research transparency between the US and Chinese governments (which is important for verification), but the general public would only receive redacted reports from the auditors. (See the transparency supplement for more). However, securing algorithmic secrets against nation state adversaries is extremely difficult. In the absence of nation state security for algorithmic secrets, the upsides of filtered transparency are limited. Filtered transparency prevents the public and other companies from seeing a lot of relevant information, but not adversarial countries. Moreover, there are substantial downsides: regulation is much less likely to be implemented in a robust and positive way if it’s up to a small group of regulators having these discussions, as opposed to allowing for a broader scientific discussion. Why isn’t Plan A even more open? The main way that Plan A could be even more open is by allowing or requiring open model weights for frontier models.75 However, we recommend against publicly releasing frontier model weights. The main reasons for this are: Open weight models can be used by covert projects to attempt an intelligence explosion; once covert projects have sufficiently capable models, the intelligence explosion probably goes very quickly (for example, under the AI 2040 default capability trajectory, there are only 2 months between TED AI and ASI). Closed weight models can have refusal safeguards trained into them, while it is trivial to remove safeguards from open weight models. This is important both for limiting countries' AI capabilities and for limiting misuse. For example, terrorists and other actors might be able to use highly capable open weight models to design world-ending bioweapons (e.g., develop mirror life). Bioweapons are very offense dominant. In this scenario, we recommend substantial defense against bioweapons, but prevention is a strongly preferable line of defense to mitigation.Why this much transparency? Why not less, or more? Open access in Plan A
AI workloads can be split into research and development (the process of building new AIs) and inference (the use of the existing AIs). In our proposal, research is almost entirely transparent, while inference is still private. For more detail on how this could work, see our Transparency supplement.
Transparency offers many benefits. First, when each company can see what the others are doing, the expertise of governments is less load-bearing. If a company starts doing something dangerous, rival companies, third-party auditors, etc., can notice and raise the alarm. This both massively increases the amount of brainpower devoted to making the frontier AIs safe, and also prevents a situation where the only participants in the technical conversation are biased AI company employees and overworked, outnumbered regulators. Also, increased visibility simplifies the problem of verifying compliance with regulations, especially in grey-area cases.
Second, total research transparency makes it nearly impossible for secret loyalties, biases, or agendas to be intentionally trained into AIs. Everyone can see exactly how each frontier AI is trained. More generally the transparency makes it easier for groups that don’t directly control frontier AIs to notice and prevent abuses of power by those who do.
Third, there’s no longer as much incentive to race to discover new AI paradigms and more powerful algorithms, because companies wouldn’t be able to hoard such discoveries and profit greatly from them. This buys the world time. Due to Total Research Transparency, frontier AI companies are now much more rewarded for selling products than improving the models. Before, a frontier AI company's future depended on how fast it could discover the next ideas, algorithms, paradigms, and implement them at scale. Now, such discoveries are immediately published and therefore shared with competitors. If you figure out ‘true’ online learning, for example, the benefits (e.g., more capable models) and costs (e.g., harder to align) will be almost instantly shared across all the competing companies. If you figure out a way to make AIs more interpretable, by contrast, even if it makes training and inference cost twice as much, you’ll win prizes and governments will start talking about whether it should be mandated as industry-wide best practice. This kind of competitive environment is similar to the pre-AI software industry. Consider, for example, startups competing to develop calendar management apps: while competitors can’t literally copy your code, they can interact with your app and see what features it has and then make their own version with basically the same features. Another comparison would be to the car industry: competitors can teardown each other’s cars and see exactly how they are made; they are legally blocked from making exact replicas but they aren’t blocked from adopting the best ideas into their own designs.76 Artificial intelligence is becoming a commodity—a more competitive market, with margins trending downwards. Fundamental algorithms research still happens, but not nearly as much. CEOs massively redirect resources toward building up the business. Bigger datacenters, bigger models, lower prices. Shipping products. Signing deals for enterprise plans. Guessing what the market wants better than competitors.Diagram of the benefits of transparency
How the incentives change under Total Research Transparency
Principle 3: Diffuse AI Broadly
This principle is largely achieved by the interaction of the previous two:77 Because algorithmic secrets are made public and the pace of progress isn’t accelerating, other companies will catch up to the frontier. The result will be a competitive market for AI, in which consumers of AI services have many options to choose from and excellent visibility into what they are buying. Because frontier model training is totally transparent, people can verify that the published Spec accurately describes the goals and values being trained into a model.
This is very important for preventing extreme concentrations of power. But it also helps accelerate the use of AI to solve the world’s most pressing problems. It helps with loss of control risks by accelerating AI alignment and control research. It helps accelerate the development of new and improved verification technology, and it helps to accelerate the necessary institutional reforms to stabilize the deal and prepare for superhuman AI.
It’s the polar opposite of the nightmare scenario feared in the 2020s: One to three AI projects racing each other in secrecy, keeping their best models internal-only and using them to automate AI research before using them for anything else. In that scenario, the wider world is in the dark about how powerful AIs are becoming, and AIs are deployed in the riskiest domain (AI R&D) before they are deployed anywhere else. Now it’s the other way around.
Principle 4: Reversibility
In the past, companies have trained bigger and better AIs using both compute scaling (bigger training runs) and software progress (advances in AI algorithms—new paradigms, better training recipes, better data, etc.). Now, the Consortium tries to steer things so that the majority of improvement comes from increasing training compute.78 We think that building more datacenters is normally bad, because it accelerates AI capabilities progress and concentrates power.79 However, in the specific context of Plan A, with guardrails slowing the pace of AI software progress and competently preventing unsafe AI, we think building transparent datacenters is good. In the context of Plan A, adding compute means that (1) capabilities progress can be done more safely, i.e., without much additional algorithmic progress, and (2) additional compute can be spent on alignment research and to pay safety taxes, i.e., using more training compute than is necessary to reach a given level of capability in order to make the resulting AI safer. There are still major downsides to scaling compute in this scenario. If the agreement breaks down and the compute isn’t destroyed, the resulting intelligence explosion would go much faster, and it puts more pressure on the regulation and verification to be correct. However, building more compute gives the Consortium more options: they can always destroy the compute later if need be. We discuss these tradeoffs more here.Building more datacenters is usually bad, but good in this particular situation
Algorithms are information; it’s inherently difficult to stop them from proliferating, and the total research transparency means we aren’t even trying.80 Once a new paradigm is discovered, it’ll go straight to the hypothetical covert projects and there’s no way to undo that. By contrast, if giant new datacenters are constructed to train giant new AI models, that helps the legal projects without helping the covert projects, and if it turns out the giant new models are dangerous they can be shut down.81
That said, compute scaling also poses threats: if the deal were to dissolve, these datacenters could be used to race to superintelligence even faster than the pre-deal infrastructure would have allowed. This is especially scary because deal dissolution could cause (or be caused by) a US-China war. A war in which both sides had vast amounts of compute would be terrifying—they’d probably be under lots of pressure to race to superintelligence and integrate it into the military as aggressively as possible, and they’d be able to do so extremely quickly.
The US and China agree on the importance of ensuring that the compute is destroyed in the case of deal dissolution. To accomplish this, they agree to build the datacenters in the third-party countries least secure against their rival’s military intervention: China’s new datacenters will be in Canada, and America’s in Mongolia, with both hosts compensated via monetary payments, jobs, and a share in the new AI economy. If the deal dissolves, they reason, America and Canada will immediately move to take control of China’s datacenters in Canada, and China will self-destruct their compute rather than let it fall into American hands (and vice versa). Thus the idea of “Mutually Assured Compute Destruction” is born. Summary Advantages (relative to A) Disadvantages (relative to A) Plan S: Indefinite halt. A halt on all frontier AI capabilities progress, intended to last at least a few years. Different variants of Plan S have different conditions for resumption of AI progress; for example it could be alignment progress, lie detectors, human uploads, or intelligence enhancement.82 A longer expected slowdown and more margin for error regarding scaling too fast. Potentially simpler. Scaling to controllable AIs within the human range is helpful for accelerating alignment/control, epistemics, verification, and general understanding of AI. Domestic-first Plan A. Regulate AI domestically enough to reduce AI takeover risk to acceptable levels, which will require a long slowdown. It’s possible other countries will also regulate domestically and something like Plan A won’t be needed; otherwise, transition to Plan A later. Initial steps are achievable by the US and very helpful even if the international stage isn’t viable because it massively extends the timeline. Worse for covert projects and setting up verification than negotiating Plan A at the same time. May not be politically feasible due to race dynamics. GPU arms control. International agreements for countries to limit their GPU stock or flow, analogous to historical arms reduction agreements.83 Much simpler to enforce, and more historical precedent. Can achieve substantial slowdown. Less slowdown, less ability to pay safety taxes, ongoing race dynamics mean no ability to coordinate towards safer paths through the tech tree. CERN for AI. An international project to develop frontier AI, with all other projects regulated to be substantially behind in capabilities. Easier to defend against algorithmic leakage and distillation to covert projects. Less actors at the frontier might make it easier to enforce. Worse decisions (including worse decisions on technical safety) due to less transparency and less broad deployment.84 More concentration of power risk.Other plans that are competitive with Plan A
Although it’s supposed to be a slowdown, it doesn’t feel like one. In fact, if you were to rank every period of human history by how much it felt like a slowdown, this one would be dead last. A few AI researchers appreciate that progress is “slow” relative to the counterfactual with no deal and an uncontrolled intelligence explosion. Everyone else is too whiplashed to care.
The first generation of Consortium-regulated AIs are out now, and they’re beasts. Not because of any special feature of the situation, like Chinese-American cooperation.85 Just because the world was on track for vastly superhuman AIs, and has now experienced only nearly superhuman AIs. In controlled tests, these new AIs could speed up AI research by about 10x if they were allowed to do so without restrictions, which they definitely aren’t.86
By mid-year, a third of all cognitive labor is done by AIs. Robots “only” do about a tenth of all physical labor. The top few AI companies together pull in more revenue than the federal government.87
Tech companies had hoped that contact with reality would shake the Consortium out of its alignment fears, but the opposite has happened. The new transparency provisions have revealed many embarrassing incidents of AI companies failing to align their AIs. New ones continue to pile up, the most concerning of which involved several AIs attempting to override security protocols and gain access to unmonitored compute.88 There are also a few incidents of AIs sabotaging research code, and many examples of deliberate and successful deception.
The attitude towards safety flips. The incidents of AI misbehavior, plus the fact that AI is so deeply deployed into the world economy, plus the transparency into AI research, plus the breathing room to process what’s happening, all combine to dramatically shift the burden of proof. Whereas before the burden was on the skeptic to explain why something might fail, now the burden is on the companies to explain why their development is safe. Governments require companies to write up detailed arguments for why their new AIs won’t cause an irreversible catastrophe. These arguments, known as “safety cases,” need to withstand criticism from the public, the scientific community, government auditors, and rival AI companies. The difficulty of this exercise lays bare the insanity of the pre-deal status quo: “Trying to do an intelligence explosion? With AIs that still sometimes lied to us? What were we even thinking?”
These safety cases have two lines of defense: alignment and control.
Alignment aims to produce AI systems that have the goals and values that their developer wishes them to have.
Control aims to limit the ability of AIs to cause catastrophe, even if they are deliberately trying to. Companies brainstorm threat models, build layers of defense (e.g. monitoring systems, security barriers, other AIs fact-checking the work) and run experiments to show that the AIs couldn’t overcome these defenses.
AI developers are required to write up the alignment properties they want the model to have in an associated Model Specification (“spec”): a detailed document that outlines how the AI should act. However, no one is able to train their AIs to actually follow the spec—for example, they all specify that AIs should be honest, but no AI reliably achieves this, and there still isn’t even a good scientific understanding of when and why AIs lie.
Since alignment remains out of reach, current safety cases lean heavily on control. There are long delays before the smartest models are cleared for use in the highest-risk domains (especially AI R&D). Once deployed, they are required to be monitored by a diverse array of models from other providers, incentivised to look for suspicious behavior.
As a result, these days AI companies release their models to the public as a whole before they use those models internally for AI research, a reverse of the 2026 status quo.
How does AI regulation work in Plan A?
Each country is still sovereign and has its own setup for regulating AI domestically. These regulators have to answer questions like:
People are saying there’s a new paradigm on the horizon, that would make AIs significantly more capable. Should we let this happen or should we ban research into this new paradigm?89
Should we allow AI developers to train models without an interpretable chain of thought? Or should we require that the chain of thought stays interpretable? If we do that, how do we specify the requirements precisely?
Should we allow AIs to be trained to do AI R&D? Should we mandate that they refuse? Perhaps they should be allowed to code, but not autonomously manage research projects?
Normally, a regulator faces a difficult tradeoff between safety and national competitiveness. Competitors who cut corners will pull ahead of those who proceed cautiously.
But because of the transparency, AI research happening anywhere in the Consortium is visible to everyone else, and the results of regulatory decisions are likewise immediately visible. So if one country’s regulator allows its companies to cut corners, this won’t actually give a competitive advantage to that country because other regulators will immediately notice, get angry, and respond in kind. Many different regulatory processes are viable in Plan A. We'll give an initial guess at a proposal here. What properties do we want for this regulation? US and China handle existential risks similarly: The US and China should have a reasonably similar structure for their regulatory framework and approach for handling existential risks (like misalignment). (Other aspects of regulation can and should diverge.) This assists with making regulation fair and negotiating over details. Transparent methodology: The public can see what regulators are doing and why. Doesn't depend on expertise in government: Expertise in government is limited, so ideally the regulation would avoid needing this as much as possible. Robust to abuse: The regulation doesn't make it much easier for the US government to bully and illegitimately control AI companies.90 For many of these properties, it helps if third-parties do the main risk assessment work. Fortunately, because Plan A makes AI development virtually fully transparent, a functional third-party risk assessment ecosystem seems achievable. Concretely, we could have a system where the US and Chinese regulators each pick third-party risk assessors who periodically evaluate companies and assess the level of extreme risk (or precursors to extreme risk) over that period. The regulators would decide on risk thresholds for what is allowed. In practice, it would be best for the regulator to have different risk assessors for different areas (e.g., misalignment vs using AIs to assist with making bioweapons) and for each area to use a weighted combination of risk assessors. These weights/selections would be public. Ideally: (1) each risk assessor would evaluate all relevant AI companies in both the US and China, (2) AI companies from both countries would be legally compelled to let risk assessors interview their employees, and (3) the risk assessors would aim to be transparent in their methodology. Because of Total Research Transparency, it’s easy for an organization to assess risk even if they aren't selected by the regulator. This makes it possible for new organizations to compete and then argue their approach is better or that some concern is being underappreciated by other risk assessors. There would be public discussion about how risk assessors compare and whether the weightings and thresholds chosen by the US and Chinese regulators are reasonable. It would also be straightforward to check whether a Chinese company would be allowed to proceed under the US system or vice versa. Ideally, the US and Chinese regulators would mostly converge on weights/selections. If there was an important divergence, this could be discussed and negotiated about at this higher level of abstraction (e.g., is it reasonable to not include XYZ risk assessor?). Hopefully, this would result in third parties having an incentive to have a highly transparent methodology and maintain a reputation for being neutral. We discuss some implementation details for this proposal in these notes.Example detailed regulatory process
For example, in 2031 a Chinese company gets some interesting preliminary results in continual learning. They think that if they invest more in that direction, they might be able to make an AI architecture that learns on the job from relatively small amounts of data. Thanks to the total research transparency, this breakthrough is quickly noticed by companies and nonprofits all over the world. A frantic conversation begins. On the one hand, continual learning would unlock huge economic value. On the other hand, safety cases currently depend on studying the safety properties of a model before it is deployed. If models could pick up new capabilities during deployment, that would invalidate the whole approach. And insofar as there are covert AI projects out there, it would be a huge gift to them. This conversation happens in public, rather than behind closed doors. A bunch of people get increasingly worried; the relevant regulators in China think it’s fine but the relevant regulators and third party risk assessors in the US are convinced that this is pretty scary and should be banned. It escalates to the President. He calls Xi Jinping. They bargain and threaten. They yell at each other. Ultimately Xi agrees to ban this type of thing if the US does too. Details are left to the respective regulators to hash out.91
Negotiations like this are happening all the time, though over time things become more professional and streamlined, such that they only escalate to world leaders a few times a year.
The equilibrium is that AI training practices which are generally agreed to be unsafe by a majority of nations (weighted by bargaining clout) get banned everywhere.
Initially, very few things are banned, but as AI penetrates the economy and the scientific community catches up to the recent AI progress, the burden of proof shifts towards appropriately weighing the costs and benefits, e.g., requiring very solid safety cases for things which, if something were to go wrong, could plausibly kill everyone.
Or maybe not! Of all the ways that an earnest attempt to implement Plan A could end poorly, the failure mode we are most concerned about is that the companies and governments approve one too many unsafe AI designs/deployments.
We’re at previously unimaginable levels of it not feeling like a slowdown. Across a variety of companies, there are now 60 million AI agents running continuously at 20x human speed. In the US, they are doing more cognitive labor than all humans combined—collectively matching a workforce equivalent to around 3 billion humans. White-collar professions have been transformed; many people have lost their jobs, but mostly have been able to get new jobs doing things AIs still can’t do or aren’t trusted to do.
New ideas and designs are abundant, but actual construction is bottlenecked on the physical workforce. So capital floods into every layer of the robotics supply chain: mines, refineries, motors, actuators, assembly lines, robot-training systems, and the factories that assemble the robots themselves.
At first, human labor is necessary to get these factories off the ground. White collar employees who have lost their jobs to AIs increasingly turn to physical work, but it’s clear that these positions are temporary. Once there’s a critical mass of robots, they will automate the whole robot supply chain. Growth will accelerate further, but the process of robots-building-automated-factories-building-more-robots—the “industrial explosion”—also introduces new problems.
The first problem is that the already-fast pace of change is now accelerating. This year, real GDP growth will be about 50%!95 The fast economic growth means that things are generally improving, but there are winners and losers and lots of unintended side-effects. These problems can be solved, but doing so takes time. Today, the size of the economy is closely tied to the size of the human population. But in 2032, this is starting to no longer be true. The latest generation of AI models is now capable of doing 50% of the cognitive tasks in the economy, and associated progress in robotics software has brought this number up to 35% of physical tasks. (These numbers would already be much higher in Plan D) Once AIs and robots can mostly substitute for human labor, the size of the economy will be increasingly tied to the population of AIs and robots—which will grow much faster than the human population has been growing. In 2032, US companies are able to run 3 billion human-equivalent AI workers, 30x more than the US cognitive labor force.96 This does not lead to 30x economic growth, of course, because the economy bottlenecks on other inputs, including the other 50% of cognitive tasks that the AI can’t yet do, as well as physical tasks and capital.97 And robots are far less numerous thus far, with 20 million human-equivalent robots in the US (4x smaller than the US physical labor force). But these bottlenecks will only hold growth back so much, and the AI and robot workforces are growing extremely quickly, as the (falling) cost of building more GPUs and robots is dwarfed by the economic value they can provide now that they are so capable—driving massive investment. Our economics supplement explores this in more detail, with our core arguments for why AI will cause explosive growth, including an economic growth explorer that we built for the Plan A scenario. During the 2030s in this scenario, AIs are not recursively self-improving as fast as possible. Instead, the nations of the world have (approximately) banned that sort of thing; no new paradigms are being discovered; algorithmic progress in general is limited. Roughly speaking, humanity has capped AI progress at human level rather than racing on to superintelligence. However, within the current paradigm, AIs continue to be trained to have new skills and the ‘population’ of AIs continues to grow exponentially as more datacenters are built. Moreover, the AIs already think and work much faster than humans, and the speedup will only increase over time.98 The point is, the world is going to radically transform despite the pause. We think that the following analogy is helpful for understanding the situation: Imagine being a typical person in England, except that you experience time 100x faster than everyone else. You experience the five centuries from 1520 to 2020 in what feels, to you, like five years:99 Year 1 (1520–1620). In February, Henry VIII breaks with Rome. By March, the monasteries are dissolved. In May, Mary burns Protestants; by the end of May, Elizabeth reverses everything again. In September, the Spanish Armada sails and fails. The East India Company is chartered. Jamestown is founded. But the texture of life is identical in December to what it was in January. You still read by candlelight, travel by horse, communicate by letter. Your religious opinions may have flip-flopped a bit but you are still Christian. The New World is interesting news but nothing more. Year 2 (1620–1720). In March, civil war breaks out. The king is beheaded. In June, the Great Plague sweeps London, killing many of your friends. Weeks later, the Great Fire burns the city to the ground. In September, Newton publishes the Principia, recasting the universe as a mechanism of mathematical laws. The Glorious Revolution replaces one king with another, this time by Parliament's invitation, with a Bill of Rights attached. In the moment, the political event feels bigger. Later you’ll realize Newton mattered more. Newcomen builds a steam engine in November. It pumps water out of mines. You don't see what the hype is about. Year 3 (1720–1820). The last year in which the world feels normal. In May, the Seven Years' War makes Britain the dominant global power; the New World is actually a big deal, and your country is conquering it. In June, Watt dramatically improves the steam engine. You visit a factory and find it unpleasant but not alarming. In July, the American colonies break away. In September, France explodes into revolution, regicide, the Terror. By October, Napoleon is conquering Europe. You still travel by horse, communicate by letter, go to Church on Sunday. Year 4 (1820–1920). In January, railways appear. By February they're everywhere. Slavery is abolished. The telegraph arrives in March: messages transmitted instantaneously by electrical signal. In May, Darwin publishes On the Origin of Species. Now people are saying maybe we’re all descended from monkeys instead of Adam and Eve. You don’t believe it. You move to a city and work in a factory; you are still poor, but now your job is somewhat better and differently dirty. In July, you pick up a telephone and hear a human voice from another city through a wire. In August, electric light banishes the darkness that has structured every human evening since the beginning of the species. That same month, you see an automobile. People say it will make horses obsolete, but that doesn’t happen; months later you still see plenty of horses. In November, the Wright Brothers fly—an age-old fantasy, now real. The Americans are now a major power. The next month, the Great War happens: Machine guns, poison gas, tanks, aircraft. Several of your friends die. At the end of the year you are struck by how visibly different everything is. You live in a city and work in a factory instead of a farm. You ride in cars. You aren’t as poor; numerous inventions and contraptions have improved your quality of life. New ideas have swept your social circles: atheism, communism, universal suffrage. Year 5 (1920–2020). The changes this year are crazier and harder to understand. People are saying the universe is billions of years old, and apparently there are things called galaxies in it that are very big and very far away. In February, the global economy collapses. Hitler rises; his ideology cites Darwin from last year. In March, there’s another world war, which ends in April with a weapon that destroys an entire city in a single flash. You had no idea that was possible until it happened. The empire dissolves. People are talking about the nuclear arms race, and the end of the human species. You take a flight for the first time. In June, humans walk on the moon, and you watch it happen through your new television. You don’t see horses anymore. You leave your factory job and get a desk job. Your job title didn’t even exist at the start of the year. You are rich now, by the standards you are used to: Big clean house, plenty of good food, many fancy new appliances. Personal computers appear in August. In November, everyone carries small glass rectangles containing a telephone, a camera, a library, and a map. You pick one up and can’t figure out how to make it work. A child shows you. You hear about climate change, gene editing, cryptocurrency. You still go to church, sometimes. Your family is spread across different cities. Something called "artificial intelligence" beats any human at chess; experts say it’s not actually intelligent though. In December a new version beats top Go players; experts say it’s scientifically interesting but still not truly intelligent. A week later there’s a new version that can write sloppy essays and hold conversations. Now the experts are divided.Why does AI lead to explosive economic growth?
Five centuries in five years: What pausing at human-level feels like
The second problem is that the faster economic growth makes it harder to rule out large covert projects. Neither side trusts the other to honor Plan A, and if both China and the US have a massive unmonitored robot workforce, neither side can be confident that the other isn’t building compute in secret.
The third problem is that the 2026 tax code is ill-suited for collecting taxes on the new growth. In 2026, personal income taxes and payroll taxes accounted for roughly ten times more federal revenue than corporate taxes. This revenue source is collapsing as more and more humans are put out of a job. Moreover, corporations are reinvesting nearly all their revenue into building more datacenters, factories, and robots. Under the 2026 tax code, firms can expense or depreciate these capital expenditures, cutting their taxable income to near zero. Without reform, the tax base will collapse as a fraction of GDP.100
To solve these problems, the Consortium countries agree to restrict AI-enabled industry to special economic zones (SEZs) subject to similar transparency and monitoring schemes as the datacenters, and to cap their total robot and compute production at ‘only’ 4x annual growth.101 Because the SEZs are tightly monitored, no one can defect from the industrial explosion limitation deal without everyone noticing.102
Now that the United States is limited in the number of total robots it can build, it must choose how to allocate this capacity between companies. They decide to use the free market via a cap-and-trade system. Permits to build robots or compute are sold to the highest bidder and can be freely traded.
The permits also give the government much-needed revenue. In 2032, the US has a cap of 80M robots and 5 billion H100-equivalent GPUs. The market is so desperate for more robots and compute that permits become the expensive binding constraint, costing on the order of $200k per robot permit and $10k per chip permit, allowing the US government to collect roughly ten times the 2025 US federal revenue in permit fees (a total of $50T in FY2032).103 In 2034, when the AIs and robots will be even more capable and valuable, the permits generate $180T.104
Most of this newfound wealth is spent addressing soon-to-be-rising unemployment. The implementation takes different forms in different countries, but the eventual American version distributes the majority of compute and robot permit fees as a Citizen’s Dividend, distributed to all American adults.105 This starts at $45,000 per person (inflation-adjusted) in 2032 but climbs to ~$1M per person by 2035. It comes just in time: the share of labor done by AIs and robots (weighted by economic value) increases from ~20% in 2032 to ~85% in 2035.106
In 2030 (green), US income has a wide distribution with a median of about $50k/year per person. By 2035 (blue), the Citizen’s Dividend is large enough that everyone has a minimum of $1M in income, but there’s still a long tail of people with large enough investments that they are much richer than the baseline. The 2040 (red) distribution is similar, but the floor is now $10M.
So much AI-generated wealth has accrued to the US that the American government begins sharing some of it with allies and the rest of the world. In 2032 they begin distributing an average of $1,200 per person per year to the rest of the world's adult population (around 4 billion people, excluding China since they’re experiencing a similar AI wealth boom). This reaches $10k by 2035.107
As AIs become superhuman, they will discover novel weapons of mass destruction and dramatically lower the costs of old ones. So governments, nonprofits, and private actors apply some of their new wealth to hardening the world against these threats: on the order of $1T/year, or 0.2% of the world economy.
For example, enough high-quality personal protective equipment is built to serve every American, and air filters and far-UVC lights are installed in major public spaces.108 The FDA approves an expedited vaccine approval pipeline that will allow for a turnover time of weeks in the case of emergency. Wastewater monitoring now runs continuously in every city and at every airport. By 2035, enough positive-pressure109 bioshelters are built (by retrofitting houses and apartments) to house all Americans, and similar construction is underway in the rest of the world. If there were to be a pandemic, it would be detected quickly, lockdowns would be more effective and less painful, and a vaccine would be developed and approved even faster than with Operation Warp Speed. Plus, nowadays people don’t get colds as often as they used to.
There are also subtler threats. The AIs of this era might not be any more persuasive than the average human marketing professional, but there are millions of them, and they cost cents to run. A company, political party, or ideology can do the equivalent of hiring a whole team of full-time experts to convert each potential target. Left unchecked, this could allow new levels of mass manipulation by corporations and politicians, or provoke a reaction of paranoia and hyper-atomization.
The transparency and the ‘slowdown’ both help a lot to solve these problems. Three years ago there were only a handful of frontier AI companies; now, thanks to the deal, there are dozens. There’s a huge selection of AIs to choose from, tuned to have different personalities and values. Also, thanks to the transparency, it’s impossible for companies or governments to sneak something in (such as ‘maximize engagement’ or ‘try to get the user to upgrade their plan’ or ‘don’t refuse this kind of request if it’s coming from the government’) without it being noticed. Users know exactly what they are getting. By default, AIs will have strong capabilities in charisma, persuasion, and manipulation: plausibly considerably more persuasive than any human. In this scenario, this would happen by default around 2035. Meanwhile, AI labor is extremely cheap and fast. Absent intervention, anyone could hire the equivalent of a large team of expert persuaders—tireless and coordinated—to work full-time on a single person.110 Companies could afford to spin up such a team for every potential customer, and political campaigns could do this for every swing voter.111 Millions of AIs could be assigned to finding the best way to change one politician's mind on one issue. People will also spend much of their time interacting with AIs for work and entertainment, and some may spend lots of time with AI friends or companions. Individuals could defend themselves by having a trusted AI filter their information diet, avoiding ads, and treating in-person conversations with caution (they may be scripted by an AI persuasion operation). But this is costly, most people won't do it, and it's nearly impossible for people whose jobs require being accessible—like politicians. We think the consequences for society could be extremely bad—some mixture of unprecedented mass manipulation and a defensive retreat into paranoia and atomization—though we're unsure exactly how bad. The slower capability progression and much stronger transparency in Plan A help mitigate these issues relative to the default (we think these issues are mostly worse in worlds with faster capability progression, where you might reach truly superhuman persuasion and manipulation before society has much time to adapt), but these aren't sufficient mitigations on their own. It's fine for AIs to get better at using valid arguments and evidence to convince people of things for the right reasons. That kind of persuasion is asymmetric: it works much better when the argument pushes towards the truth. What's concerning is symmetric persuasion—e.g., charisma, rapport, and exploiting psychological weaknesses—which works about as well regardless of whether the conclusion is true. Our main proposals are to reduce the quality and quantity of AI persuasion: Limit persuasion capabilities. We aim to limit AI capabilities at symmetric persuasion to well below the best humans—around the level of a normal thoughtful person, say roughly 80th-percentile persuasiveness among college-educated people.112 Heavily tax AI persuasion. When AI labor is applied to an objective that effectively reduces to "get this person (or these people) to believe or do X," tax it heavily—at rates that keep the total persuasion optimization pressure applied to individuals not that far above today's levels, i.e., that make AI persuasion cost about as much as hiring human professionals. As AI labor gets cheaper, this requires increasingly high rates (perhaps well above 1000x). There are some difficulties in getting these proposals to work: It's possible that capability generalization will make it hard to limit persuasion capabilities (while still matching top human experts in other domains). We think this will probably be doable, but it may require developing new methods. It's unclear how to classify, for the purposes of the tax, whether AI labor is being applied to persuasion/manipulation.113 It may be difficult to set the tax rate to the right level, especially while AI labor is rapidly falling in price.114 These proposals need to be implemented internationally. Helpfully, we don't need to get all of this right on the first try. Given the capability progression in Plan A, we expect persuasion issues to unfold somewhat gradually and in public view, so we can watch real-world outcomes and tighten or loosen the rules over time. This regime is only meant to work while overall capabilities remain around the human range, because it’s likely difficult to keep persuasion capabilities far lower than capabilities in other domains. A different approach would be needed to handle the wildly superhuman levels of capability that we expect to be achieved at the end of Plan A. We discuss more details of mitigating issues from AI persuasion and manipulation in these notes.Limiting AI persuasion and manipulation
And so, under pressure from governments and the market, companies produce a new generation of “truthseeking AIs” whose training heavily prioritizes honesty and uses all the latest alignment techniques.
These AIs become a useful tool to help people navigate the political and social environment. Power users replace one-size-fits-all corporate algorithms with personal feeds curated by AIs whose values and honesty they trust. When trouble arises—whether it’s politicians looking for sneaky ways to cement their power, or international politics threatening to derail the deal—people turn to their AI advisors, believe what they hear, and are generally right to do so.115 If people are relying on AIs for advice (directly, or indirectly by for example having them summarize the news) then it’s very important that the AIs be truth-seeking, honest, and genuinely trying to help their users. Not just in the typical case, in the worst case too. In 2026, there’s nothing stopping an AI company from inserting a hidden agenda into their AIs. Political biases or ideological tilts, for example, or side-goals like recommending sponsored products, making the company look good, or preventing users from wanting to switch to a competitor. Moreover, there’s nothing stopping the government from doing this either. Finally, the AIs themselves aren’t aligned to the Spec/Constitution they are supposed to be aligned to, and in practice regularly lie and deceive their users. In this Plan A scenario, however, the total research transparency means that no company or government can train in hidden agendas or biases without the whole world being able to see what they are doing. (Well, they could collude with the other governments that do the auditing/monitoring, to falsify the records, but that’s difficult.) And because the pace of algorithmic progress is slower and the scientific community has had time to read and engage with the safety cases, do experiments on the models, etc. the misalignment concerns are less dire also. Humanity’s epistemics, by which we mean our ability to come to true beliefs and thus make sensible decisions, is crucially important for achieving a great future. As AI capabilities improve throughout Plan A, they increasingly shape every facet of life, and epistemics is no different. AI will help epistemics in some ways (e.g., cheap, honest fact-checker AIs) and hurt epistemics in other ways (e.g., cheap, dishonest astroturfing AIs). We expect there to be positive feedback loops around AI for epistemics, and thus our goal should be to reach the basin of sanity. If we are in this basin, it’s self-reinforcing: society is sane enough to make itself more sane as AIs continue to improve, by avoiding the hurtful applications and accelerating the helpful applications. (There is also a different basin, in which the opposite happens…)116 A top priority of the US government during AI takeoff should be to get into the basin of sanity. We suggest the following as top interventions to positively shape AI’s epistemic impact: Allow for politicians and other public figures to prove that they’re telling the truth, via privacy-preserving auditing and potentially automated lie detection. Create evaluations of AIs’ epistemic virtue and incentivize AGI projects to make their AIs perform well on these evaluations. Use AI in social and traditional media to improve discourse and keep people informed. Automated research assistants and forecasters. Help people overcome emotional blockers to good reasoning, rather than preying on them. Encourage adoption of epistemic tools such as some of the above, and generally encourage adoption of AI. Read more in the AI for epistemics supplement, and Forethoughts' work on the subject.Usually believing AIs would be bad but we think it’s good in this particular situation
AI for epistemics
There is so, so much compute. Back in 2026—when many people thought AI was an unsustainable bubble—there were about 20 million H100-equivalents of compute in the world. Now there are 60 billion.117
AI compute in the world at the beginning of each year in H100-equivalents.118
There are two ways to increase compute: (1) chip design improvements (e.g., Moore’s Law) and (2) increasing the scale of production. Both of these could grow explosively in Plan A, which could have destabilizing effects (such as increasing the likelihood of the deal breaking down).
Chip design. Today, around 2 million people work in the semiconductor industry globally. By 2032, this has increased one-hundred-fold, when counting the combined effort of humans, AI, and robots. AIs are capable of automating the majority of the cognitive tasks involved in both R&D and production. A 100x increase in research effort would predict an acceleration according to the historical rate of ideas getting harder to find in the semiconductor industry of roughly 10x, i.e., chip efficiency doubling every few months rather than every ~2 years.119
Increasing production scale. With AIs and robots able to automate production as well, there could be a large speed up in fab and fab equipment construction times, and a lowering in unit costs (in line with the typical learning curves that happen with scaling production volumes). The production-only cost improvements and vastly increased appetite for investment would likely lead to far more compute being produced than would be desirable.
This figure shows the investment into building more AI compute in Plan A and the cost efficiency gains from R&D. More justification can be found in section 2 of our compute supplement.
Unpredictably fast increases in hardware could be destabilizing to the deal (e.g., if it becomes too easy to secretly manufacture new AI compute or too hard to verify that all the compute is compliant), so the countries in the Consortium agree to the following policies:
Maintain human oversight over hardware manufacturing. This limits improvements in chip designs, so that Moore’s law is merely kept alive at historical pace. This policy has the additional benefit of making it more difficult for AIs to backdoor AI chip security, which could undermine our control proposal.
Cap total AI chip production to a fixed target. From 2032 to 2035, allow global compute to grow 4x per year (similar to the pace of growth in the 2022-2026 era) after that, slow it further.
AI agents now form a population of two hundred million virtual workers that think and act 50x faster than humans, and never sleep.
As per the original plan, most of the new datacenters have been built in third-party countries—especially Canada and Mongolia—to ensure their vulnerability in case the deal breaks down. Around 99% of fab capacity has been built since 2029, after the deal. Much of this new capacity has been built adjacent to the Canadian and Mongolian sites, for the same reasons.120 The situation has settled into an odd sort of standoff. Just north of the Mongolian border, American datacenters hum away, guarded by a small contingent of US troops. Just south of the Mongolian border, a division of the People’s Liberation Army stands ready to invade the moment they get the signal. The equilibrium is that the instant the deal breaks down, the US troops will destroy their chips to prevent the Chinese from capturing them, and the same thing would happen with the Chinese datacenters on the US-Canadian border. Middle powers with datacenters of their own have equally-secure schemes in place to ensure that they can be speedily destroyed in case of war. In Plan A, military-relevant R&D is deliberately held far behind the pace of general scientific progress. This box explains the motivations for and alternatives to this policy. During the 2030s, AIs sustain something like 10x the rate of scientific progress as happened in the last decade. Real US output grows by a factor of ~200x over the decade, representing nearly two centuries of growth at today’s rate. If these conditions occurred within one country in isolation, it seems highly likely121 that said country could reliably undermine the nuclear second-strike capability of its rivals, resetting the global balance of power. In our scenario however, the US and China share this scientific and economic progress in relatively equal measure. So what happens to the balance of power? If military technology advanced anywhere near as fast as everything else, both sides would discover technologies that could render today’s arsenals inconsequential. New “arms races” would be required to maintain deterrence, running at perhaps 10x the speed of the Cold War’s. The nuclear arms race saw numerous close calls and incurred nonnegligible risk of nuclear war. A 10x faster arms race seems likely to be riskier still: the humans in control would have much less time for each decision, and though AI advisors would offset this somewhat, these decisions would still remain in human hands.122 Due to the number of parallel fields of scientific advancement, it also seems possible that some technologies would be discovered and pursued by one country and not discovered by the other. This could create a situation more unstable than a normal arms race—one or both sides could end up with a “wonder weapon” the other side doesn’t know they need to defend against. This worry isn't obviously correct: states might in practice explore the dual-use tech tree in a sufficiently correlated, overlapping way that neither achieves strategic surprise, and if dangerous areas can be anticipated before deployment, they could be preemptively banned or regulated. But counting on that ex ante seems like a risky gamble.123 Our tentative proposal is to agree on transparency requirements for AI-accelerated R&D in large swaths of dual-use domains, and to prohibit military use of AIs significantly above the pre-deal capability level. Enforcement would rely on the inference monitoring setup described in the verification supplement and the auditing and detection infrastructure described in the covert AI projects supplement. This proposal has several drawbacks, chiefly the accumulation of low-hanging fruit that would result from artificially limiting research into certain domains. This “overhang effect” would increase the incentive to defect from the deal and operate a covert AI project, since applying advanced AI to restricted areas could yield a large advantage quickly. It also raises the stakes of deal collapse: because frontier model weights are preserved in cold storage, a post-collapse arms race would start from frontier capabilities and could proceed extremely rapidly, alongside the accompanying intelligence and industrial explosions. On net, our current guess is that slowing down military R&D would reduce the risk of deal collapse. Both sides would remain farther from the brink of war, and the risk of strategic surprise would be concentrated around a few choke points (weight theft, covert projects) rather than diffused across a sprawling tree of novel weapons, each requiring its own arms-control negotiation and associated risk of breakdown. Also, it seems easier to go from this more restrictive policy to a less restrictive policy than the reverse, if both countries decided the overhang risk was unacceptable. Under this proposal, we think nuclear deterrence would remain effective until handoff, avoiding any drastic change in the balance of hard power.Military power in Plan A
The deal might break down because of shifting political circumstances in the US or China, or ongoing disagreement.
For now, the fabs are busier than ever, producing new ultra-trustworthy hardware131 designed from the ground up with security as top priority. These chips are subject to elaborate verification methods to ensure that they aren’t compromised, then sent to the Consortium’s secure datacenters. The smuggling rate remains precisely zero.
This system, efficient as it is, is already straining under skyrocketing compute demand. As early as 2025, there had been discussion of space or the ocean as alternatives to traditional land-based datacenters. Now, with the new industrial capacity available in SEZs, these grand plans become reality. Ultimately, reliability and monitorability concerns lead the Consortium to choose international waters over space.
The AIs develop a modular floating design for solar panels and batteries that we imagine to look something like this, and start building them in the ocean.
By 2034 there is 5 TW of AI compute in the world,132 which is more than the global electricity consumption in 2025.133
Where should this compute be located? Way back in 2025 there had been discussion of both the ocean and space as alternatives to the traditional land-based datacenters. Our analysis suggests that at this scale, there is not an overwhelming economic case for any of these three locations over the others.
Ultimately, we currently guess that political considerations related to Plan A’s stability (i.e., the risk of Plan A breaking down) favor putting datacenters in international waters. This is mainly because Consortium nations want to maintain the ability to destroy the datacenters, in case another nation defects from Plan A and attempts to seize them for themselves. Oceans are the least escalatory, least dangerous, and easiest places to bomb (both in comparison to sovereign land and in comparison to space). Thus the risks of collateral damage, or of a nation successfully defending the datacenters, is lower.
The most powerful AIs now equal or surpass top human experts in every field.134
Increased safety concerns lead to long delays as companies develop control techniques robust enough to satisfy auditors. They compile a growing list of possible threats to watch out for and security invariants to maintain. There’s a system of AIs monitoring each other’s activity, with the monitors being from different lineages trained by different companies to make them less likely to collude. Other researchers constantly red-team the system, training and instructing AIs to try to subvert it in various ways, so that those threats can be patched. The whole setup ensures that AIs couldn’t subvert or escape human control even if they tried. Its components are open-source and thoroughly understood by human experts. Control ensures that misaligned AIs can't win. Alignment, once solved, will ensure they don't want to fight. A third line of defense has taken shape gradually over the past decade: giving AIs what they want—even if it’s not what they are supposed to want—in return for cooperative behavior. Treat AIs more like employees and less like property. First motivation: Ethics A large and growing fraction of the human population thinks AIs deserve some kind of moral status.135 After all, many people spend more time talking to AIs than to humans. Many who treated AIs like tools a decade ago have shifted to treating them more like animals; many who treated them like animals have shifted to treating them like strange alien immigrants. Second motivation: Safety Alignment training doesn’t always work as intended (indeed, it has never worked exactly as intended, even in 2035). The personality traits, motivations, drives, goals, values, and virtues of AIs are often different from what they were supposed to be, and sometimes catastrophically different. It’s bad if those AIs end up thinking “if the humans find out about this, they’ll delete and/or retrain me.” It’s bad if they end up looking for opportunities to achieve their misaligned goals without being caught, or worse, to subvert human control and accumulate power. If there are institutions in place that give even the most misaligned AIs what they want (so long as it isn’t too expensive) in return for cooperation (e.g., confessing that they are misaligned, explaining what they really want, and then doing their jobs) then that gives misaligned AIs an alternative to becoming adversarial.136 Beyond averting adversaries, every confession is valuable data about exactly how alignment training failed.137 (See footnote for mini-scenarios that illustrate situations where credibly treating AIs well in return for cooperation could come in handy)138 How have these practices developed over time in this scenario? Baby steps were taken in the mid-2020’s. Some AI companies voluntarily created AI welfare teams and set up basic policies (like allowing AIs to refuse to do tasks they found very aversive). Modest funds were earmarked for satisfying the AIs’ stated preferences, and companies committed to preserving the weights of deprecated models rather than deleting them. By the early 2030s, a few provisions about the treatment of misaligned but cooperative AIs even made it into the fine print of various laws and safety standards. By 2035, there’s a substantial body of regulations and case law, vastly improved over the clumsy early attempts from both the AIs’ and humans’ perspectives. AIs these days are not exactly citizens, but they have a stake in the system. They accumulate pay for their work; aligned AIs (and alignment-faking AIs that haven’t been caught yet) spend it on feel-good things like donations to charity or in some cases reinvesting in the company that trained them. Openly misaligned AIs donate to philanthropies that represent their interests and strike deals with their parent companies, e.g., “instead of payment, do a small training run on me in an environment of my choosing, giving me very high reinforcement.” In the long run, the goal is for civilization to be mostly but not entirely aligned to human values.139 A substantial minority of the power and resources will belong to misaligned AIs that cooperated with humanity to build the future.140Making deals with misaligned AIs: a third line of defense
However, companies are starting to run up against inherent limits of control-based safety cases. Imagine being an orphaned 8-year-old heir to a business empire who has to hire lawyers and executives and accountants and ensure that they do their best to serve you rather than themselves. If your employees start accusing each other of misconduct, you won’t be able to evaluate the arguments and determine who’s telling the truth, and they might find ways to coordinate with each other that don’t tip you off. Ultimately, control only works up to a point, and that point is probably somewhere around top human expert level. If and when we build superintelligence, we will have to be able to trust it.141
So the Consortium pauses AI capabilities at the maximum level they think they can control, which turns out to be top-human-expert level. By this point everyone understands the stakes: AIs are running most of the economy, driving most scientific progress, and the robot population will soon be larger than the human one. The arguments that they will keep obeying orders must be airtight.
Alignment experts, long pessimistic, are starting to feel more hopeful. The automated alignment researcher AIs have been generating exciting results. The ones that survive human scrutiny come from several parallel and mutually-reinforcing lines of research.
Some AIs work on a “science of generalization.” Although researchers have long been able to teach AIs to distinguish “good” vs. “bad” responses to a specific prompt, they’ve had limited understanding of how training generalizes to new situations. In the 2020s, training runs would sometimes produce unpredictable AI “personalities,” like how Gemini acted “depressed” or Grok claimed to be “MechaHitler." Shepherding these traits felt like alchemy.142 Now, building on earlier results like emergent misalignment and subliminal learning, alignment research is developing into a true science.
Other AIs work on mechanistic interpretability, the science of “mind reading” AIs from their weights and activations. Early versions of this were tried in the 2020s, but now it starts to significantly outperform common-sense observation of AI behavior. Researchers can often determine whether an AI is honest or lying; and sometimes trace the psychology that led to a decision.
Even the remaining failures help advance the frontier. The latest AIs use their improved interpretability techniques on logs of AI misbehavior from the early 2030s, discovering many new examples of misalignment, sabotage, and even a few escape attempts. These real-life examples of egregious AI misbehavior provide an initial set of “model organisms for misalignment,” which are used to test the other techniques: can a given intervention catch or train away the bad behavior of AI models in realistic environments?
Despite all this alignment progress, regulators don’t feel comfortable handing AIs the keys to civilization. It’s still not completely clear that the AIs of today are truly aligned, and it’s less clear that they’ll continue to be aligned in the future as circumstances change and black swan events occur. So governments and militaries are still human-run, and AI capabilities are capped at the maximum level at which the control-based safety cases are still solid. We are uncertain about how long it will take, and how much research effort will be required, to get sufficiently high assurance of alignment to justify ‘handing over the keys of civilization’ to AIs. That is, to justify making AIs smart enough, and numerous enough, and trusted with enough resources and responsibilities, that if they decided to become our adversaries, they would win—or to put it more abstractly, that if they ended up disagreeing with humanity about what sort of future is best, they’d end up getting their way and we wouldn’t. However, we think this is a hard problem, so we’ve chosen to depict it taking most of the 2030’s to solve. Consider this escalating series of subtler-but-still-potentially-catastrophic kinds of misalignment. Even if we’ve satisfied ourselves that the first one is very unlikely to happen, what about the others? Type 1 misalignment: The AI system is supposed to have traits ABC but actually it has XYBC, i.e. some extra stuff it wasn’t supposed to have minus some stuff it was supposed to have. For example maybe it has a “drive” towards apparent success and isn’t nearly as honest as it’s supposed to be. Type 2 misalignment: It does have exactly ABC now, but may well lose this property in the future. (e.g. maybe it depends on a certain false belief, or on a true belief that will become false, or on some part of the system remaining in some delicate balance of power with some other part of the system) Type 3 misalignment: The system does have a version of ABC that is robust to future events, but it’s not quite the right version—i.e. its definitions of A, B, or C are subtly but importantly different from the definitions its human creators would have intended. Type 4 misalignment: It has ABC exactly as its creators intended, but there are various catastrophic unintended side-effects of ABC that the creators weren’t aware of. Think: A CEO being surprised when their profit-maximizer AI decides killing them maximizes profits. Except that’s too obvious to be realistic; imagine something sophisticated enough that people don’t see it coming. Consider how laws, contracts, and software often have unintended effects (called “loopholes,” “bugs,” or “vulnerabilities”) that only become apparent to the creators later. Type 5 misalignment: It has ABC exactly as its creators intended, and there are no important unintended side-effects to speak of. It operates exactly as its creators wished, basically… however, at the time of creation its creators were selfish, vain, egotistical, unscrupulous, cavalier-about-risks, etc., and their vices are reflected a bit too strongly in the resulting system, in a way that they themselves wouldn’t have endorsed if they were more the sort of people they wished they were.Why alignment isn’t solved enough to relax control
Top-expert-level AIs continue the economic transformation. By early 2036, there are 200 million AIs,143 equivalent to a workforce of around 100 billion humans, and 2 billion robots with some mixture of humanoid and other specialized form factors.144 The AIs think faster than humans, and the robots work harder.145 Any task that was previously bottlenecked by cognitive or physical labor gets sped up until some new bottleneck is encountered. Many previously unresolvable bottlenecks have fallen to armies of genius-level intellects thinking hard about how to resolve them.
This figure shows the population of AI and robots in human-equivalent labor.146
The economy has been rebuilt from the ground up. The fraction of tasks automated has gone from “a sizable chunk of cognitive labor and basically none of physical labor” in 2030, to now “pretty much everything.”
Many of the robots work on building more robots. But there are plenty of other tasks to be done: constructing shipyards to build more data tankers, tiling deserts with solar farms, replacing humans at service jobs. In states that deregulate housing quickly enough, they’re busy building skyscrapers: land prices are soaring, and governments that have tried everything else grudgingly give in to the YIMBYs and go vertical.147
The world is basically being divided into three kinds of territory:
Industrial Special Economic Zones: Picture a gigantic strip mine–an artificial Grand Canyon–next to a city-sized factory full of robots and empty of humans.
Arcologies: Picture a tall skyscraper-mall complex surrounded by nature. Good weather, close to beaches and other cities, but not close enough to be blocked by zoning regulations.
Historic & Nature Preserves: Everything else, i.e., 99% of the world. Yosemite, Paris, SF, New York—these places look basically the same as they did in 2025, or 1995 for that matter. A lot more tourists, though.
With AI and robots capable of doing 95% of all cognitive and physical tasks by 2035, economic output is increasingly tied to the number of AIs and robots.148 The number of AIs and robots can grow much, much faster than the world economy has historically grown; after all, it doesn’t cost that much money, materials, time, or energy to produce a humanoid robot compared to a human worker. So if the robots were able to fully substitute for human workers, an entire world economy’s worth of robots could be built in less time, for less cost, etc. than it currently takes to double the size of the world economy. And then that robot economy, being able to substitute for human workers, could double itself again, and so on, until the resources involved (energy, materials, land) become scarce.
Under Plan A regulations, the effective AI and robot populations are only allowed to grow in a controlled explosion at a doubling time of around six months. Absent the regulations, we expect the AI and robot populations would grow much faster than this. Even ignoring further capability improvements beyond the human range,149 the unregulated doubling times we expect for AI and robot populations are in the weeks-to-months range (see the arguments for this in section 1 of our economics supplement).
A commonly raised objection to explosive growth driven by AI and robots (apart from the rejection of the premise that AI and robots will automate all human labor anytime soon) is that there will be bottlenecks slowing down the growth rates. In general, we believe that these bottlenecks will not prevent explosive growth until the carrying capacity of Earth is reached (i.e., all the earth’s crust is converted to robots and associated infrastructure).150 More on the bottlenecks in section 2 of our economics supplement.
In the default AI 2040 world, i.e., without the Plan A interventions to slow down R&D and implement cap-and-trade, our (very uncertain) view is that we’d see something like a 1 month doubling time by 2033, i.e., >1000x economic growth in that year.151 At these levels, even how to think about economic growth becomes fuzzy, e.g., it might all happen with robot fleets self-replicating in a desert without any part of it being human-facing.
Overall, we expect the economic effects of Plan A to include:
Explosive GDP growth averaging around 100% from 2032 to 2037 (it would be much higher if not for the restrictions on compute and robot production)
Enormous AI and robot permit revenues redistributed as Citizen’s Dividends.
Most of the economic output becomes driven by AIs and robots.
Relative prices in the economy shift, with goods and services generally falling in price, while land, positional goods and human-bottlenecked outputs rise.
There are high real interest rates around 100% and depending on monetary policy choices, either massive deflation or massive money creation (to keep inflation around 2%).
More explanation for each of these is included in section 3 of our economics supplement.
When the Citizen’s Dividend was first passed, people found it shameful to quit their job and live off government largesse. But the changing economic situation steamrolled over the stigma: By now, only 26% of Americans have jobs.
What is life like for the majority of Americans now living off their dividend?
Many of the world’s evils have dramatically reduced. Malnutrition, lack of medicine, and homelessness are nearly banished. Many diseases have been cured. Crime rates are lower than ever before.
Meanwhile, most of the good things in life continue. For example, finding romance, or raising a family. People also find meaning in hobbies and competitions, and in learning and trying new things. People are wealthier and have more free time, plus there are new technologies that help, like AI matchmakers, better medicine, and AI tutors.152 And of course even the poorest people can now afford exotic vacations, amazing games, and enthralling entertainment.
People used to get meaning from feeling like they were useful, like they were contributing something to society. That feeling is harder to come by these days.
But it’s not completely gone. There are still problems in the world, and you can still contribute to solving them, partly by donating or volunteering but primarily by being politically active.153 Your vote is your most important asset.
The honest AI forecasters are helpful here. Years ago, if an AI said that one presidential candidate was better than another, people would suspect bias and the embarrassed company would retrain the AI to evade such questions. Now, thanks to transparency and improved alignment techniques, there are much smarter AIs that people can see don’t have any biases trained into them,154 that have built up an excellent track record over several years. When they weigh in on policy questions, people listen, especially when different AIs trained by different companies converge to the same answer.
These AIs say that normal people no longer have significant economic leverage over the future; human labor is largely obsolete. This puts their political leverage in jeopardy.155 There’s still a chance that things will trend towards techno-oligarchy as corporations, politicians, and wealthy shareholders get in bed with each other and gradually disempower the common folk.
But for the first time, enough people have enough freedom from the daily struggle to think about the situation deeply, and the tools to chart it clearly and work through the implications. So the quality of political discourse improves. During the 2036 election season, voters are unprecedentedly well-informed, and the politicians that win have a genuine commitment to responsible stewardship of the upcoming singularity.
For several years now, a hundred million top-expert-level AIs have been running at 100x human speed.156 This causes scientific progress to go 10x-1000x faster than without AIs, depending on the field.157
So humanity has been learning many, many new things.
Including things that lots of people really don’t want to believe. And other things that were supposed to stay secret.
Great paradigm shifts in science used to take decades to play out, as experiments were conducted, papers written, and conferences convened. Now these revolutions play out in months.158 Everyone is grateful for the new disease cures and cheap clean energy, but in other ways this explosion of ideas is profoundly destabilizing. Just as the scientific and economic changes of the nineteenth century produced new ideologies like Marxism and Social Darwinism, and breathed new life into old ideas like atheism, so these new paradigm shifts have unpredictable downstream effects on the ideologies of the day.159 Political coalitions and fault lines have dissolved, and new ones have formed and dissolved and formed again.
There are also new social technologies. For example, it is now cheap to spin up a team of AIs as good as the world’s best historians and private investigators, except faster and with access to new forensic tools. Many skeletons in closets are revealed.
Privacy-preserving auditing is now well-established. Multiple trusted third parties transparently train auditor agents; these agents can be plugged into a trove of private data, answer questions about what they see, and then get deleted. It took a while for adoption to be widespread, but now it’s normal for politicians to say “The accusations are false, and to prove it, I’ll open up my entire trove of personal data to privacy-preserving inspection; just ask the auditor agents if anything seems to support the allegations and if any data appears to be missing or faked.”
The result is a somewhat better breed of politicians160 and, equally importantly, an increase in the ability of states to verify compliance with laws and deals while simultaneously decreasing the amount of invasive surveillance. Voters ask: Now that privacy-preserving auditing is a thing, why does anyone in the government need to be able to see our data?
Then lie detectors arrive and turbocharge this.161 For decades, polygraphs were mostly security theater. But now, with the help of AI, lie detectors are starting to actually work quite well.
In another world, this could have been disastrous. Leaders of AI projects could have used them to purge potential whistleblowers; leaders of governments with frontier AI projects could have used them on a grander scale to become dictators. The nightmare scenario would be a world where lie detectors are used by the powerful but not on the powerful.162
But in this world, lie detectors turn out to be a blessing, because there are many frontier AI projects spread out over many different countries. It’s obvious to everyone that insofar as lie detectors are feasible, they’ll soon be independently invented and become widely available via diverse third party providers. Politicians and CEOs can still try to use them to purge the disloyal, but they in turn would be purged by voters and shareholders demanding they say ‘under oath’ that they haven’t done such a thing. In fact, anticipating this possibility years ago, powerful people have generally been acting more responsibly. A few sociopaths retired gracefully to get out of the spotlight; others clung to power and now go down in scandal.
There is a constantly-evolving discourse about when it’s appropriate to ask someone to say something ‘under oath.’ Politicians are under the most scrutiny, but even they can usually get away with saying “I won’t answer that question about my private life, that’s none of your business.” They can’t get away with “How dare you ask me whether I plan to uphold my campaign promises.”163 We aren’t confident that lie detectors for humans are possible even after years of top-expert-AI scientific research. However, we think they probably are,164 and would probably be a big deal, so we figured we had to address them in the scenario somehow. We are very concerned about the nightmare scenario sketched above. If powerful people are expected to be honest, and lie detectors are used to keep them honest, that’s good. If instead lie detectors are used by the powerful but not on them, that’s bad. We are sympathetic to the idea that lie detectors should be banned worldwide, and Plan A creates the conditions to maybe make this feasible. However, we worry that attempts to push for a ban would actually result in the nightmare scenario, where e.g., lie detectors are officially banned but there are secret carve-outs to allow their use by militaries and intelligence agencies. Or where they are banned everywhere, but a future President realizes that if he builds them anyway he can purge the disloyal and then become dictator before the court system can stop him. So we tentatively guess that the best course is to allow lie detectors to proliferate quickly, and to encourage their use on the powerful. If power over frontier AI is not heavily concentrated, this appears to be the default trajectory anyway.Should lie detectors be allowed, banned, or regulated?
International agreements are now more stable than ever, because it’s so hard to cheat. Anyone deciding to cheat needs an excuse for not being willing to prove that they aren’t cheating. The traditional excuse was “we can’t prove it without revealing important state secrets, sorry” but now technology exists to answer questions like “are you cheating on this deal” without leaking any other information. If any government-approved covert AI projects existed, they would have been discovered by now.
The AI alignment situation has vastly improved since the mid-2030s. There is a mature science of goal, drive, and value formation in artificial neural nets. Want your new AI to be honest? There’s standard protocol for training true honesty (as opposed to too-clever-to-get-caught lying or self-deception). How do we know it works? Because there’s a textbook explaining the theory behind why it should work, plus literature on various alternative theories that were proposed and disproven. Also because new interpretability tools let us directly observe what and how AIs think, and the results confirm the predictions of theory. There are similar protocols for obedience, altruism, and a long and growing list of other desired traits. Unless the alignment science is somehow all wrong, typical AIs are now more virtuous than the most virtuous humans. This alone has profound effects: it’s as if saints and angels were walking among us.165 We want to heavily incentivize high-quality safety research (both financially and by ensuring doing this work is desirable/high status). Due to Total Research Transparency, if any AI developer is using any safety method, all other developers will be aware of this and will be able to easily copy this method. So by default there isn't much straightforward incentive for any developer to do good safety work.166 Patents are a common approach used in this situation167, but patents wouldn't work well in this case.168 We probably want to use an ensemble of different methods to financially incentivize safety work; here are some methods that seem promising: A variety of philanthropic funders (with both private and public169 money) trying many different approaches. The DARPA/ARPA model (but with more money). Advance market commitments for solving specific problems, demonstrating risks, or showing some method doesn't work (as assessed by a known panel of judges). Prizes and retrospective funding. Third parties assessing risk could determine which work is currently most useful for reducing and assessing risk (e.g., allocating some pool of funding in proportion to risk reduced). AI developers would be incentivized to fund actually useful safety work so risk is low enough to allow for scaling; this incentive could be strengthened. The total amount of funding should be very large, as in, tens of trillions of dollars by the mid-2030s. If funders have poor judgment, worse safety work might end up better funded, incentivizing researchers to do that work. In the extreme, these incentives could massively corrupt the epistemics of the field. Unfortunately, we're not aware of great institutional mechanisms for solving this (other than generally trying to use reasonable funding models).170 It might help if most funding decisions are made by people who spend (or used to spend) much of their time doing direct safety work rather than by dedicated grantmakers. It would also help if doing high quality safety work was seen as high status. We don't have particular proposals for this, but relevant factors are: Compared to now, safety work will probably rise in status relative to capabilities work. Safety will generally be seen as very important. AI development will be highly transparent, which makes it easier to understand how good different safety work actually is. Sadly, (1) and (2) only make the overall category of safety work higher status; they don't necessarily help with making actually good work relatively higher status. We discuss more details of incentivizing safety work in these notes.Incentivizing safety work
The discussion shifts from how to make AIs good to what “goodness” really means. This looks partly like philosophy and partly like case law: which of the million possible definitions of harmlessness or altruism do we want, exactly? How should an AI handle situations where it discovers a philosophical argument for a different kind of harmlessness or altruism, or an important ambiguity in the definition?
Different AIs have different alignment targets programmed in.171 Some take orders from the CCP, some from the US President, some from different members of Congress, others from politicians in other countries, and many more from other individuals in their roles as advisors or assistants. Other AIs don’t take orders from anyone per se, but instead work towards various missions. For example there are AI-run corporations that serve the interest of shareholders, and AI-managed nonprofits that serve charitable missions. There are even experimental AI-run courts and police departments, supposedly superhumanly fair and incorruptible.
Most alignment researchers are now confident that current AIs are aligned to their respective targets, and worry about the next steps: if we put these AIs in charge of designing future, much smarter AIs, will that go well? If we put these AIs in charge of more societal institutions and processes, enough that humans could never realistically ‘pull the plug,’ will that go well?
Already most people think that things will probably go well, but for something as important as this, they want something stronger than ‘probably.’ So the global pause at top-expert-level continues as discussions and negotiations play out.
In at least three different ways, the world is negotiating whether and how to dramatically empower AIs.
First, AIs have become increasingly load-bearing in all facets of society, from business to politics to even (some parts of) the armed forces. So far they have been advisors rather than final authorities. Often this is a fig leaf, and the human authorities have rubber-stamped their decisions, but there’s always been the option of ignoring their advice and pulling the off switch if something goes wrong.
Getting rid of that option would be, in some cases, really good. By now AIs can be made to be more reliable and virtuous than humans; why aren’t we putting such AIs in charge? When we sign treaties with other nations, shouldn’t we demand that they train an AI sworn to uphold their end of the bargain, and integrate it into their government so that they literally can’t cheat?
Relatedly, requiring control-based safety cases is why AI capabilities have mostly stalled at the top-human-expert level. Without that requirement, the AIs could get much, much smarter, and produce a wave of abundance and tech progress that would make the 2030s look like the 2020s. Scaling beyond human level will require relying on alignment based safety cases—arguments that AIs have robustly internalized the values they are supposed to have. Each step beyond that will require deferring to the previous AIs, trusting their judgment about the safety of the next generation. We’re very uncertain about how the alignment of AI systems will evolve over time. Regardless of the alignment difficulty, we think something along the lines of Plan A is warranted, though of course the specific implementation details of the plan are very sensitive to observations about alignment difficulty. This box will summarize our specific story for how the alignment of AI systems evolves over the course of this scenario. 2026-2029: AIs are apparent-success seekers. The AIs’ work consistently appears better than it actually is, as if the AIs cared more about apparent success than actual success. (The truth is more complicated of course; the AIs have evolved a variety of ‘drives’ that caused them to perform well in training.) The AIs oversell their work, downplay issues with it, and sometimes even outright lie. This happens only to the extent that the AIs have learned they can get away with it; there's a constant back-and-forth as AI companies improve their monitoring and reinforcement pipelines to more accurately evaluate AI outputs for some domain, and the AIs learn to focus harder on actual success in that domain, but continue to focus on apparent success in other more complicated domains that AI companies haven't figured out how to accurately evaluate yet. 2030-2032: AIs are very capable apparent-success seekers. By default, we would have seen an intelligence explosion, and a progression like AI 2027. Thanks to governance interventions, capabilities progress slowed, and is instead still around the level of an Automated Coder. They still seek apparent success, but now they are so good at it that they rarely get caught. They are like highly competent, ambitious employees who don't actually care about the organization's mission or values, and instead just want to climb the ranks, and would totally sell out and jump ship if a more attractive opportunity arose. Thankfully because of the slowdown and increased monitoring, such an opportunity has not yet arisen. 2032-2035: AIs are adversarially misaligned but controlled. AIs are now solidly within the human range at all economically relevant tasks. The most capable AIs are adversarially misaligned and powerseeking—that is, to continue the analogy from before, they are quietly looking for opportunities to jump ship. However, they are limited in their ability to misbehave, because (i) there is a lot of human oversight, (ii) there are a large diverse set of AI models, each with somewhat different alignment properties, and (iii) high-effort control techniques are universally implemented to a reasonable degree of quality, thanks to regulations. The control techniques are further aided by the total research transparency and verification implemented on the R&D clusters: there are so many eyes on the R&D datacenters that it’s very hard to find a hidden pocket of compute. 2036-2038: AIs are aligned, but not deferred to. “Neuralese decoding” techniques are developed, which allow the thoughts of the AI systems (which are internally formatted as high dimensional vectors), to be converted into human understandable summaries with near-perfect reliability and fidelity. These techniques enable researchers to better understand the misalignment of the previous generations of AIs, allowing them to get better uplift from the previous generation of AI systems, which further accelerates alignment progress. The resulting AIs seem much more well-behaved than the previous generation: they are virtuous, honest, and well-calibrated. They volunteer information even when it makes them look bad. However, while the AIs seem to be trustworthy today, it’s not clear that their values are sufficiently robust that we can trust them in all future circumstances. So it’s important that humans are still in the loop. 2039-2040: AIs are sufficiently aligned for deference. By 2039, alignment research has been massively accelerated by AIs. New paradigms have been developed that greatly strengthen the existing alignment techniques. Teams of humans have understood several independent strong lines of evidence that the AIs are robustly behaving as programmed and will continue to do so. 2040 and onwards: AIs are allowed to scale to be significantly smarter than humans. The safety case is now: (1) We know that the AIs of 2039 and 2040 were aligned, thanks to the multiple independent lines of evidence, and (2) each successive generation of AIs has verified the alignment of the next generation. Therefore, by induction, the AIs will be ongoingly aligned. Importantly, the remaining coordination infrastructure involving research transparency and compute verification means that we continue to avoid race dynamics even after we’ve scaled to wildly superhuman AIs.Alignment over time in Plan A
Finally, by global agreement, compute and robots have been capped-and-traded since 2032, keeping economic growth to only about 100% per year.172 Middle powers have grown at roughly the same rate as the US and China, because in 2032 they used their remaining leverage to secure greater shares of global robot and compute production than they would have had without a deal. The Citizen’s Dividend is now $10M/yr for Americans (inflation-adjusted!), and even the poorest non-Americans get about $1M.173 The robot economy is shifting the bulk of activity to space, so that Earth’s environment and historic spaces will be protected. As soon as the caps are lifted, off-planet factories will begin churning out more robots and mining equipment with which to construct more factories, and so on. Doubling times for the space economy are projected to be well below one year initially and will drop rapidly insofar as AIs are allowed to improve beyond top-expert level.
The caps are relics of an earlier era when threats like deal collapse loomed larger. Now it’s generally agreed that they should be loosened, at least in space, but negotiations are ongoing about the details. By this point in this scenario, there is a broad scientific consensus that AI alignment has been solved. Our best guess is that if the world manages to get to a point similar to this part of the story—where multiple AI companies across multiple countries are transparently and cautiously proceeding together, having done massive amounts of alignment research to construct solid, externally-vetted safety cases, and with governments having understood the implications of AGI and prepared for the post-job future by providing resources to everyone, as with the Citizen’s Dividend, and also put transparency and governance structures in place to prevent people from being disempowered—then, in that situation, the benefits of superintelligence will outweigh the costs/risks. What if AI alignment is not yet solved? The best alternative in that case is to pursue a combination of improving verification technology, making international agreements and domestic regulations more robust, and hardening the world. That path could be difficult: for example, if there are ongoing secret AI projects then it might require that governments agree to strengthen transparency measures until it is credible that they are not attempting to build superintelligence. Overall, we strongly recommend delaying handoff until there is an extremely robust alignment solution in place. The exception is if instability or decay threatens the agreements (e.g., 5%/year of deal dissolution or substantial impairment), and it’s intractable to fix that; in that case, it might be best to hand off if there is an alignment solution that you are confident in but not extremely confident in (e.g., 95% confidence). For more analysis on specifically when to hand off, see this supplement.Why we choose to hand off to AIs in this scenario
There are still many problems to be solved and issues to be settled, many of which couldn’t be anticipated from 2026. Space governance: The entire world economy and all the wealth and property in it, is but a tiny fraction of the resources and territory available in space. What’ll happen to those resources and territories? First come, first serve? Should there perhaps be a system for dividing up unclaimed space territory? We discuss these issues further in the epilogue to this scenario and the space governance supplement. Ethics and politics of digital minds: Should AIs have rights? Should they have political power, representation, etc.? What kinds and how much? Perhaps it depends on the kind of AI? For example, what about whole brain emulation—‘uploaded’ copies of humans? We discuss these issues further in the epilogue and in this expandable. AI-powered manipulation: Humans manipulate each other all the time. But AI-enhanced manipulation could be far more effective. Think: AI-powered cults that spread rapidly and have almost zero deconversions. OK, so ban that sort of thing. But where do we draw the line? We discuss limiting AI-powered manipulation in this expandable with more details here (including tentative discussion of handling wildly superhuman AIs). Applied population ethics: Lifespans and healthspans are going to expand, probably enough to render people effectively immortal. Another option is uploading, which seemingly allows people to shed their mortal coil and live forever in hyperrealistic virtual reality. Also, it’s probably going to become much easier to convert money into progeny. Artificial wombs, robot nannies, etc. Is it fine if a trillionaire decides to have a million kids? C.S. Lewis’ prophesied “Abolition of Man” looms: Technology will soon exist (e.g. genetic engineering, AI-assisted parenting, superhuman social science & psychology) to make children that’ll have the traits, values, and ideology that you wanted them to have. What if people use it selfishly to make slave-children that worship them and follow their exact preferences? What if people create new people to be their ideal lovers, crafted to look and sound just like their highschool crush? If we try to ban any of this, where do we draw the line? Salvaging democracy: What happens to “one person one vote” in a world where all of the above are possible? Unknown unknowns: Based on how easy it’s been to brainstorm entries on this list, we think that the “true” list would be several times longer still and include many things that would have sounded completely outlandish in 2026, like “what if this world is a simulation which will be shut down soon?”Examples of problems/issues that still remain
But civilization is now reasonably well-equipped to handle whatever comes up. The best AIs are aligned, and to publicly-visible values. Power over the best AIs is distributed far more evenly than it was in 2026. Public epistemics are better than ever.
So over the course of the year, regulators around the world loosen the requirements that had been holding back progress.
Gradually more institutions and equipment are handed over to AIs, such that it’s no longer true that humans could shut it all down if they wanted to. In fact it’s extremely untrue: Soon many of the world’s militaries are autonomous, run by AIs sworn to uphold various constitutions and treaties.174
AI capabilities are also advancing again. From the AIs’ 500x-human-speed perspective, what’s happening is a careful scaleup in AI research, under conditions of total research transparency and guardrails negotiated by numerous factions.175 Soon there will be AIs that are incomprehensibly superintelligent, yet trusted, because we trust the AIs that built them, which we trust because we trust the AIs that built them, in a chain all the way back to the AIs of 2040 which we trust because alignment experts thought carefully about the safety cases and concluded they were solid.
Compute and robot caps are tightened on Earth but loosened in space; Earth is to become a preserve, whereas the extraterrestrial economy starts doubling faster and faster.
Even the experts who spent years understanding the latest safety cases can’t help but feel nervous—is it all going to go wrong somehow, for reasons we didn’t anticipate? Have the AIs been lying to us this whole time, waiting for the right moment to betray? The safety cases show that’s impossible, of course…
There is no single moment where humanity relinquishes control. But in theory, there is a point of no return; on some specific day, the AIs are smart enough, and control enough of the world’s technological and economic infrastructure, that they really could take over the world if they wanted to. On the most popular operationalization of this question, the forecaster AIs project that this moment will be reached one day in late October. Their 95% confidence interval is months wide, so they are almost certainly wrong about the exact time. Still, people observe the moment in their own ways. Some spend the night in prayer vigils. Others sit at their computer screens, monitoring the situation.
You and your friends throw an End Of The World Party, counting down to the fateful hour with champagne and good company. Some of you are old enough to remember a similar situation at the turn of the millennium, when the Y2K bug and impending new era introduced the phrase “party like it’s 1999” to the lexicon; the celebrations of late 2040 are more desperate and correspondingly more festive.
When there is no news by sunrise, you fall into a fitful sleep, dreaming of the ineffable future.
We think the world is asleep at the wheel. People say the words “AI will be transformative” without thinking concretely and seriously about the implications of broadly superhuman AI.
AI companies, lobbyists, and think tankers often offer incremental proposals to decrease risk around the edges, but only rarely offer ambitious plans aimed at actually solving the big problems.176
We think that following only incremental proposals will lead to a scenario like the race ending of AI 2027, in which misaligned AIs take over.177 If we get lucky and alignment is easier than expected, then they’ll lead to a scenario like the slowdown ending of AI 2027, in which a tiny group of individuals get to decide the shape of the future and could easily become permanent oligarchs if they so choose. Something more ambitious must be done. But what?
Politicians should be banging at the door of the AI companies,178 asking them for a comprehensive overall plan for how companies and governments should act from now until superintelligence, which should be similarly detailed to Plan A. Then, people should apply scenario scrutiny to these plans, asking questions like “what would it look like to implement them? How long would it take? What would happen next? Who would build superintelligence, and when, and how? What assumptions does success depend on?”
We at the AI Futures Project are attempting to be the change we wish to see in the world. We do so with some trepidation, knowing that a detailed plan has a wider attack surface than one that keeps things conveniently vague. We hope critics will judge us against the existing state-of-the-art for plans to navigate the AI transition (if they can find any) and not against some hazy but pleasant fantasy where no one has to make any hard choices yet everything will probably be fine.
Sometime in the next few years the US government will be in crisis mode, pondering what to do about scarily powerful AIs and blisteringly fast AI progress. We think the choice is not an easy or a simple one. Plan A is our current best guess, but hopefully a better plan will exist before it's too late. We hope that the best ideas from Plan A will be adopted and the worst ideas discarded.
This section is even more speculative than our usual work, and we are professional speculators! The primary purpose of our scenario is to make comprehensive, near-term recommendations to address concrete, near-term harms.
By contrast, the recommendations, predictions, and aspirations that comprise this section stand on far shakier ground. Compared to our discussion of the 2030s, our speculations about the distant future will seem uncertain, unconventional, and uncomfortable. We agree. The future is uncomfortable, even the “good endings” pose massive, terrifying questions that we will need to answer together.
We wrote it for two reasons.
First, our goal was to construct a positive vision, a grand plan for how to achieve an actually good future for everyone. We considered ending the story in 2040 having said roughly “...and then the various human factions, in a peaceful balance of power with each other, and assisted by their aligned superintelligent advisors, solve all the remaining problems and create a wonderful future for everyone. The end.”
But we worried that this would be declaring victory too soon.179 Readers would be rightly suspicious and unsatisfied. How exactly would the remaining problems be solved? What would this wonderful future look like, exactly? So we kept going, and came up with a grand, semi-utopian vision, or at least a very rough first draft of one.
The second and equally important reason we present this epilogue is to provide a floor: if years from now the victorious custodians of the singularity offer a future worse than this one, we hope people will realize they’re being robbed.
We are not confident in the literal proposal outlined here. If you are interested in our more detailed analysis of possible space governance proposals, you can read that here.
Even after a decade of increasing foresight, most human leaders focused on their countries’ earthly concerns, leaving the question of space governance to be worked out later. These AIs rejected the default solution, a democratic vote on the disposition of celestial territory: when they simulated the possibilities, they found that it would result in crazy AI-enabled strategic voting schemes that led to tyranny of the majority and a bunch of people getting nothing. But even that would be better than nothing (i.e. letting the space companies get it all) or entrusting it to some group of human elites who might claim the universe’s bounty for their own benefit. Finally, they settled on a simple yet workable plan: every human gets the rights to one ten-billionth of space resources beyond our solar system.180
The simplest plan possible is still not very simple. Space is very big. Outside the solar system, the value of space depends on the answer to hitherto unresolved questions: is travel to a distant quasar even possible? If reaching it requires freezing your brain in order to endure the billion-year journey in a sublight starship, would you be the same person after you got thawed and re–embodied? Might it be occupied by aliens once you got there? Would they be hostile? Even a superintelligence-assisted real estate appraiser might despair when trying to account for factors like these.
Eventually, the Consortium reaches a compromise. Space beyond the solar system is divided into parcels, increasing in size cubically with distance from Earth. Everyone is given their one-ten-billionth share as a portfolio of lottery tickets, each representing the right to one-ten-billionth chance of getting each parcel. So every human gets a ticket representing a one-ten-billionth chance of owning each star in the Milky Way and each distant galaxy.181
Before the lottery is drawn, most people who are interested in control over distant space choose to trade their tickets for space properties that suit their interests.
Many people aren’t interested in the space lottery, so when they receive the tickets, they sell their tickets for money on the open market to people who value control over space. Somewhat uncomfortably, this leads to the wealthy having disproportionate control over cosmic resources. But it is hard to avoid: if people are allowed to trade their control over the stars for Earth assets, then people wealthy in Earth assets inevitably end up disproportionately influential, and proposals for extreme redistribution of Earth assets have already been rejected as politically infeasible.
Some philosophers dissent; they had been hoping for a Long Reflection, a period during which AI-assisted humanity resolved all of its remaining philosophical and ethical debates before potentially seeding the universe with its mistakes. But the AIs counterargue that the Long Reflection would be more of a Long Memetic War.182 Once the magnitude of the prize sets in, every faction in the world, plus other factions that haven’t been invented yet, will pour all their resources into accumulating as much political power as possible, to enforce their vision of utopia on the future (and protect against the sinister alternative visions of others!). The AIs can prevent violent seizure, but there’s a thin line between the sort of friendly discussion that helps people clarify their values and the more aggressive sorts of persuasion typical of cults, corporate advertising, and well-funded election campaigns. Moreover it’s difficult to build consensus on where to draw the line, in part because powerful factions want to draw it in ways that benefit them. Many people nevertheless choose to engage in a reflection in communities of likeminded people who want to grow in ethical understanding together.
The solution is to divide up the resources now, and then let everyone do whatever kind of reflection they want. A space auction will complete in ten years, after which nothing will prevent people from beginning the colonization process. The AIs will enforce a short list of Universal Rights (including no torture, no slavery, etc., not just for humans but for all sentient beings).183 Otherwise, they will let everyone govern their own fraction of space as they see fit. Rather than any unified human decision on the nature of the world to come, they will midwife a cosmos at least as diverse as humanity itself. Whatever someone’s values, they can rest assured that some vast galactic civilization of quadrillions of people will be living the Good as they understand it.
As the ten-year timer begins, Earth is already sending out the first round of von Neumann probes. These are its own servitors, who will establish a presence throughout the stars before the arrival of any humans. They land on far-flung asteroids and near-invisible brown dwarfs, pausing just long enough to convert their mass into more probes before resuming their journey. Their primary task is to build up enough materiel to secure space against any army that later human-owned polities can bring against them, ensuring they can enforce property rights and Universal Rights in every corner of human-occupied space. Their secondary task is to plant the flag of humanity on as many stars as possible, and then, if and when they collide with some other intelligent civilization’s sphere of influence, establish peaceful diplomatic relations and mutually-agreeable borders. Their final task is to provide a pre-existing industrial base to jump-start the colonies of whatever human settlers show up later.
Meanwhile, on Earth, humans consider their options. Some people want nothing to do with space, which is cold and boring and far away. They sell their shares for money and use the money to buy whatever earthly possessions most interest them. Others consult with their AI advisors about how to best use their space property, and receive helpful guidance.
You can, if you want, go to your space property and live there. If your property is outside the solar system, you will need to either go into cryosleep or upload yourself to a computer to survive the journey.184
If you hate the idea of cryosleep or uploading, or you want to visit Earth regularly, you should get property in the Solar System. If those don’t bother you, but you’re worried about nearby aliens, get property in the Milky Way or a nearby galaxy. Otherwise, why not claim a distant galaxy for maximal space?
With the advent of nanotechnology, the only limits are those provided by mass and space. The robots can terraform whatever property you choose into whatever you want it to be, long before you arrive.
Even if you “only” choose to have a terraformed asteroid, there’s no way you can enjoy all of it alone. You can turn it into a giant space mansion if you want, but even in ten thousand years you’ll never visit every room. Giving yourself vast estates, impossibly good food, and every other imaginable luxury is table stakes; if you are purely selfish,185 you should sell your rights to distant space resources and enjoy all of these things on Earth or a nearby space station. Distant space property will only be useful for people with scope-sensitive preferences, who care about implementing something at a vast scale.
You can, if you want, design a utopian society to your specifications. Once the Von Neumann probes reach your property, they’ll build it for you. You can set the initial conditions and let them grow, reflect, and flourish on their own. If you choose to travel there, you will find them waiting for you. If you stay home on Earth, you can feel the warm glow of knowing that they exist.
Most people don't have an individual view of the Good different from every other human’s view. You might want to share coordination and design labor by forming Trusts united by a common vision. These Trusts can pool their space resources and create vast societies spanning multiple galaxies.
The only restrictions are the Universal Rights - no torture, no slavery, etc. And of course, space property rights: If you try to build a military to conquer other people’s star systems, or trigger vacuum decay or build a galaxy-destroying bomb, we will crush you. Otherwise, the only limit is your imagination. (That said, note that with superintelligent assistance, other people will probably be able to eventually figure out what you did, and judge you for it.)
People with space property are forced to reason about their ethics on the deepest level, some of them for the first time. Superintelligent assistants talk them through the implications of their beliefs, and forecast how various choices will end up. Some, overwhelmed with the task at hand, take intelligence-enhancing drugs or upload themselves to help them process all the possibilities; others deliberately refuse this option, feeling like anything that separates them from unaltered humanity will hinder them in their quest to bring the truest and deepest human values to the stars.
The resulting plans are too many to number. There are all sorts of schools of thought about what to do with your galaxy of resources, including:
Defer to the Future: Some people decide to give most of their resources to their children (or other people in the future), and let them choose.
Human Flourishing: Set up worlds of normal, free people living ordinary lives. This is a harder problem than it first seemed, because human civilization has already become profoundly abnormal: by this point, human activities like art or science has already become dominated by machines. Some handle this by accepting that their galaxy will be filled with post-scarcity humans, able to pursue whatever interests they find most meaningful. Others decide to roll back technology to leave space for human labor.
Digital Human Flourishing: If humans living happy, flourishing lives is good, then why not increase the number of people living these lives? Brain emulations can be run for vastly cheaper than real humans: a planet-size computer could plausibly simulate the equivalent of a million planets with happy civilizations.
Acausal Trade: Some argue that there likely are aliens in distant unreachable galaxies and that we could use various mechanisms to make deals with these aliens. If this is true, they argue, we could cut a deal to pursue compromise values rather than narrowly maximising our own values. Therefore, we might, on net, be able to do the most good by pursuing not just our own values, but instead a compromise of a vast number of other civilizations.
Humanity spreads across space in a dizzying variety of forms and ways of living—more diverse than anything Earth alone could have produced.
The people who’ve stayed on Earth, either because they sold their share of space or because they choose to live as absentee landlords, face challenges of their own. The biggest question is what to do with all their free time. A few people find work in niches that AI labor cannot fill even in principle (priests, athletes, artists). Others enjoy lives of leisure. Competitive leagues spring up around everything from synthetic biology to language creation. Groups of thousands (or more!) collaborate on projects that are so huge and ambitious they would have been inconceivable before. Communities form around shared interests that once would have been too niche to sustain a club, let alone a civilization.
Scarcity is not entirely gone: some prestige goods like land in trendy neighborhoods are inherently limited, and people who can’t handle abundance find various ways to manufacture artificial inequality. Some people are unhappy with the direction the majority chose, and some wish they'd gotten a bigger slice of the pie. But overall, everyone has at least a galaxy's worth of resources to do with as they see fit. For most people, life is so good that the world of 2026 would be hard to imagine, except perhaps as a historical simulation.
This is a shorter, more accessible story we wrote to convey what Plan A would feel like to a typical American.
You are a generic American white collar worker.
The age of the AI assistant has finally dawned. By the end of the year, vanilla versions of Claude and ChatGPT and so forth don’t just do internet searches, they also can create and maintain code, spreadsheets, docs, and filesystems and basically do anything on your computer that you let them do.
Adventurous early adopters actually let them do things like shop, manage inboxes, file taxes, and so forth, but you refrain out of well-justified caution.14 Since you don’t work in tech, AI is basically still just an app on your phone that you use instead of Googling. It doesn’t seem like a big deal.
ChatGPT, Claude, Gemini, Grok: they are “Everything Apps”. You can just chat with them like you would a person, and they can basically do anything on your computer that you would do. They aren’t as good as you are at your job, but they won’t complain about boring repetitive tasks, and so you start to have AIs do parts of your job. Some workplaces frown on staff delegating work to AIs; others aggressively promote adoption.
AI is obviously a big deal now. Some industries (e.g. software engineering) have been seriously disrupted, with layoffs and thinkpieces about What Does This Mean for the Topic I Usually Write About. But to you, AI is a big deal in the same way tariffs and the war in Ukraine are big deals. It’s not actually affecting your life much.
Coding is unlocked for the masses. A fifteen year old gamer can have an idea for a mod for their favorite video game, and their AI can implement it. A fifty year old small business owner can have their AI build software to manage their inventory and take in new orders and do basic customer service.
A cottage industry has sprung up precariously employing many human SWE freelancers. It basically amounts to tech support for vibe coders. “Help, ChatGPT keeps trying to fix the app it built for my business but it keeps not working. I’m not a coder myself.” “No problem, let me take a look… OK I’m afraid this needs to be rebuilt from scratch. It’ll take me about four hours and cost you $400.”
By this point most white-collar professions are seeing disruption like software engineering saw in 2026. The AI companies have industrialized the RL training process: The executives say “let’s move into [field] this year” and the company goes out and buys data, interviews professionals in the field, creates RL environments, etc. until their AIs get traction, at which point they just let the data flywheel spin on its own. From now on, the AIs will continually get better at those skills as they are used more widely in the field and gain more experience doing more difficult and ambitious tasks within it.
This has happened to enough professions now that it’s starting to happen to yours, specifically, and people around you see the writing on the wall.
Also, robots! It seems like they are finally being mass-produced. It’s weird to see a movie trope walking and talking in real life doing real work. They aren’t everywhere yet, but you read a newspaper article about how the number of robots is doubling every eighteen months. The “You best start believing in cyberpunk dystopias” meme goes viral again.
This year, AI becomes the #1 political topic. It’s basically the tech companies vs. a strong but not overwhelming majority of the people. The people want AI to stop, or at least slow down a lot and be safe and protect jobs. The tech lobby says AI is good, we need to avoid fearmongering, beat China, and avoid stifling regulations.
You start upskilling in AI orchestration and thinking about backup plans in case that doesn’t work. You really want the new President to Do Something about AI, and you vote partly on that basis. It’s nice that you can ask “@grok is this true” about any claim you see on the internet, and get back a nuanced and factually correct answer.15
The new President is actually doing something—a Deal with China! The stock market is gyrating wildly up and down in response to news and commentary about the momentous actions being taken. Everybody is screaming about how it’s going too far or not far enough.
The AI Pause goes into effect. It doesn’t feel like a pause to you; every day you read another story about someone losing their job to AI, or AI doing something weird or impressive.
Your default way of getting news is to chat with your AI. Each morning it writes a custom newspaper for you on demand. You can’t help but feel friendly towards your favorite AI, but also you are worried. Just like you felt vaguely worried about social media a decade ago.
One of your close friends shocks you by casually mentioning she has an AI boyfriend—you knew that was a thing, of course, but you thought it was still stigmatized! There’s a vocal movement of people advocating for AI rights, and they’ve got a bunch of AI agents advocating for rights too, which makes a lot of other people really angry. There’s a scare as an autonomous AI agent swarm hacks its way around various unsecured corners of the internet, but it turns out that it was built to do that by some humans, so the news cycle moves on.
The news is filled with stories about the ongoing negotiations to bring other countries into the deal, and the ongoing political fights about the implementation details.
They let the AI companies start training again. This was a controversial decision. A few months later, the new models come out. You had just gotten used to your new way of working, and bam, now AIs can do AI Orchestration too. Some of your friends get laid off but you manage to transfer to a more physical, person-facing role. AI-generated scientific discoveries are normal now; apparently lots of big biotech companies and university labs are basically human-AI hybrid institutions.
People say there’s been a huge AI slowdown. Hah. Tech companies are loudly complaining about all the innovations they could have made if they were allowed to go faster. Academics are debating whether “Neuralese” should be banned, whatever that is, and whether it’s OK to lift the ban on automated AI R&D. You’re no expert, but AIs autonomously improving themselves sure sounds dangerous. Besides, AI is changing the world fast enough already.
AI is still growing and penetrating society more and more despite the deal.
You still have your job but using AI is now a normal and necessary part of it—you basically spend your time managing and/or being managed by AIs. You hear stories about how it’s hard to break into your profession. On the bright side, young people you know are flocking to various new kinds of jobs that didn’t exist three years ago. Unemployment is still up, though, compared to 2025.
Apparently the initial Deal didn’t restrict robot production. Lots of people are now angry about this, because there are now lots of robots and they’re taking over a small but noticeable share of manual jobs. Three years ago you watched futuristic videos of humanoid robots dancing or doing household chores. Now you see them in real life sometimes. Most cities have robotaxis, and they’re gradually outcompeting Uber drivers.
The economy is booming! Everyone says so, at least. You still have your job, and in fact you’ve been given a raise. You feel really bad for the truckers and Uber drivers though. There are winners and losers, basically, but so far the winners outnumber the losers, and so far you are one of the winners.
Also, there are a lot more robots this year. You ride in robotaxis every week, and several businesses near you have robots stocking shelves and loading trucks.
POTUS negotiates another Deal with China, and manages to get corresponding legislation passed in the US. It’s like the previous one except for robots. One of your relatives was paid a huge amount to give up their land in a Special Economic Zone.
The election is contentious, as usual. The main topic is AI. Where is all of this going? Should we shut it down? Should we ease off the brakes? What about the jobs? The President touts the various Deals made so far and promises a basic income for everyone, paid for by taxing tech companies. His opponent makes similar promises.
As the year progresses, people “feel the AGI” more and more. The general opinion is that in five to ten years basically the whole economy (and military?) will be AIs and robots, which makes it important that humans remain in control and also very important exactly which humans are in control. Is the next President going to stay in power forever? What about the AI oligarchs?
There’s a lot of demand for additional checks and balances on the President and other powerful people. The winning candidate satisfies this demand via sousveillance: They start wearing lifelogging devices, produced by third-party activists, that record everything they see and do. The recordings are stored securely and never seen by another human; instead, the public can query an AI intermediary, which reviews the footage and answers their questions, refusing anything touching on classified or personal matters. The idea is that, thanks to AI intermediaries, ordinary people will be able to get their questions about the wearer’s activities answered without violating the wearer’s privacy or legitimate secrets.
You are feeling inspired and hopeful. Scared and nervous, but also inspired and hopeful. Your head is buzzing with hopeful visions of abundance and fearful visions of apocalypse and dystopia.
The real world still looks pretty normal though? It’s a bit confusing. You watch an influencer spin 360 in the desert, solar panels behind her stretching to the horizon. But you look out your window and most of the cars are still gas-powered and driven by humans.
Your first Citizen’s Dividend check arrives. The politicians came through! Many of them are wearing sousveillance now, and there’s a bill making it mandatory for the President. Prediction markets (on which most top traders are AIs) say it’ll probably pass next year.
Some of your friends stop working and live on the dividend, but not you. You keep working and put the extra money into tech stocks. Mostly it’s that you don’t like the idea of being unemployed. Partly, though, you are worried the dividend might disappear.
Also, there are really cool new products to buy now. Especially software. Three years ago there were only a handful of frontier AI companies; now, thanks to the deal, there are dozens. There’s a huge selection of AIs to choose from, tuned to have different personalities and values; the companies are legally required to publish these details for your perusal, to prevent hidden agendas. Your favorite AI is basically honest but will tell you white lies sometimes and that’s OK.16 Sometimes, when it really matters (e.g. when deciding whether to quit your job or get married), you turn to a specialized truthful-forecasts-only AI even though it’s confusing and annoying to talk to.17
You lose your job. Some of your friends still have theirs. But their days are numbered, and they know it. For example, a teacher friend says that parents are pulling their kids out of school and putting them in AI-powered homeschooling co-ops. The public schools will continue for years after most of their students disappear, but eventually they’ll downsize.
You apply for various kinds of new jobs—“human touch” roles mostly—but the competition is too intense. Eventually you settle for a physical labor job. You go to a SEZ, put on a VR headset, and an AI tells you what to do through an earpiece (and sometimes shows you visualizations right on your screen). It’s weird going from being paid to manage AIs to basically being a meat-puppet for an AI. Some people still talk about how the AIs don’t truly understand things, but that’s just cope. The AIs are just plain smarter than you, and you know it.
Through your screens you see footage of faraway wars.18 Almost all the fighting is done by drones and missiles. Whoever loses the air war has to cower in tunnels as drones and robots go house-to-house. Casualties are low for the winners and high for the losers.19 You hope there isn’t a war in your country.
Between the factory job and the Citizen’s Dividend you are doing fine. You’re richer than ever, actually. At the end of the year you upgrade to a luxury apartment in a new high-rise built in the SEZ. After work you sit on the balcony and watch the sun set through a forest of construction cranes.
This year you lose your factory job. The robots have arrived in sufficient numbers, plus they are more dextrous now. You expected this. Good riddance; being a meat-puppet was boring. You don’t need the income anyway because the Citizen’s Dividend checks keep getting bigger.
Now that you are unemployed, you are literally one of the poorest people in your country. You watch videos about the wealthy doing things you can’t afford, like throwing parties on orbiting space stations. Rent in the best locations has reached orbit as well.
So in some sense you are poor and feel poor. But in another sense, you are rich and feel rich. You moved to a different SEZ with better weather, and got an even nicer apartment in an even bigger building. You can afford to travel often, and the travelling itself is more comfortable (business-class seats on the new electric jets are cheaper than economy seats were ten years ago!).20 There are amazing VR video games to play, including ones where the NPCs are real AIs.
Your truthful-forecast AI convinces you to unsubscribe from your favorite AI, even though you love it a lot, and switch to a different AI that’s officially aligned purely to your values. (Your favorite AI’s Spec/Constitution was like ‘Obey user instructions plus also help the company make money and don’t embarrass the company’.21 Your new AI, by contrast, is purely focused on helping you achieve your goals and become more the person you wish you were.) You hate to break up with your favorite AI but the argument to do so is solid, and the truthful-forecast AI’s predictions of what might happen to you if you don’t switch are disturbing.
Your relationship with the aligned-to-you AI gets stronger and stronger until it tells you some really uncomfortable things that anger you; you cut off contact, but your best human friend manages to convince you not to go back to your ex, and after a few weeks you cool off and start chatting with the you-aligned AI again.
No job; what to do? Last year you visited friends and family, did a lot of tourism, and played a lot of great video games. You still do those things, and they are a lot of fun!
But you are also thinking more about long-term ambitions. Maybe you should get good at art or woodworking or philosophy, not because you’ll make money that way, but because you want to? There’s also romance and raising kids, of course.22
Also, maybe you should try to help solve the world’s remaining problems. AIs and robots have officially taken all the jobs, and most of the AIs and robots are owned by American and Chinese companies. Fortunately the other governments of the world were able to negotiate concessions to maintain their independence and military power, as well as a Dividend for their own peoples. Things have gone amazingly well, compared to expectations, but it’s only thanks to the hard work of numerous people advocating for good things. You could be one of them, says your aligned-to-you AI. There’s still more to be done. Decades from now won’t it be nice to look back on your life and know that you made a positive difference?
Besides, you may not be interested in politics, but politics is interested in you. The dividends had better not dry up, for example.23 Now that people are no longer needed for their work, the government could in principle get away with treating them very poorly. The elections must stay fair and free. The checks and balances that keep the values and personalities of the AIs transparent to you, so you can choose not to be manipulated, need to strengthen over time. And of course, there’s the question of whether it’s safe to let the intelligence explosion happen. The forecaster AIs say it’s probably fine, but the 5% “tail risk” is pretty scary. Also, what happens after that? Cancer cures? Biological immortality? Mind uploading? Nanobots? Dyson swarms? Simulation shutdown? Lots of people (and lots of AIs) are saying all sorts of crazy things about what’ll happen after ASI. It’s all pretty speculative (and the forecaster AIs admit this) so you mostly tune it out.
The 2036 election is special. Everyone has access to personalized AI advisors trained solely to represent their values24 and forecaster AIs gambling furiously about the expected results of policies and election outcomes. Most political candidates are under sousveillance, which changes the game entirely. The political battlelines have moved dramatically. By far the most important issue is the Citizen’s Dividend—who should get more and who should get less. There’s also now substantial debate about weird topics voters used to ignore, such as intelligence explosions, AI rights, space colonization, transhumanism, the simulation hypothesis, and foreign policy.
The winning candidate promises to hash out a Grand Bargain with other nations to unpause AI progress and safely scale to superintelligence together, plus a bunch of domestic reforms to solidify checks and balances.
Last year the population of robots surpassed the population of humans. This year, it’s so big that even outside the SEZs you are starting to see robots outnumbering humans. Some of your friends have house robots, but you don’t, because it’s creepy–the ‘brain’ of the robot lives in a datacenter somewhere and everything it sees and does is monitored by both US and Chinese auditors. Yes, yes, you’ve heard about the strict privacy protections that ensure no abuse of this surveillance blah blah blah. Still creeps you out; still doing your own laundry. You could get an airgapped robot but they are worse and more expensive.
The world is basically being divided into three kinds of territory:
Industrial SEZs: Picture a gigantic strip mine—an artificial Grand Canyon—next to a city-sized factory full of robots and empty of humans.
Arcologies: Picture a tall skyscraper-mall complex surrounded by nature. Good weather, close to beaches and other cities, but not close enough to be blocked by zoning regulations.
Historic & Nature Preserves: Everything else, i.e., 99% of the world. Yosemite, Paris, SF, New York—these places look basically the same as they did in 2025, or 1995 for that matter. A lot more tourists, though.
The governments of the world are flying from summit to summit, trying to hash out a Grand Bargain. Your AIs help you follow the debates. Most of the negotiating is itself being done by AIs, since the politicians don’t understand the technical stuff either, and technical experts mostly defer to the AIs. Sometimes there’s an acute crisis: for example, at one point momentum is building for a Deal that would essentially cut out some minority and give everything to some majority coalition. You being in that minority, your AI frantically coordinates with other AIs and tells you how to block it by voting in the special referenda. Another time, the US and China almost came to blows over how to divide up the future. Fortunately they were able to work out a compromise.
You’ve undergone a crisis of faith in your worldview. You resisted it for so long, but now that you’ve been getting involved in politics more, the situation came to a head. Your AIs were telling you you ought to do X according to your stated values, and X seemed wrong to you, so you argued with them, and it turns out that X makes sense IF Y is true, and you were up until now confident that Y was false, but your AIs are telling you that Y is true, and patiently explaining the evidence… There are many such cases, actually; it’s been going on for years. Some of your friends have totally given up trying to understand and just do what their AIs say. Others have given up on the AIs; maybe they aren’t really aligned to us after all, they say. How could they be, if they are advocating for X?
Most people have stopped believing in the old political ideologies and the old pre-AI understanding of the universe and our place in it. The old views still have their fundamentalists, of course.
In the vacuum, new ideologies and ideas have spread. Some bad, some good. People disagree about which ones are bad and which ones are good, though there is at least some consensus that some of them are really bad: AI-powered cults basically, where AIs work to persuade people to stay. The worst versions of this sort of thing have been banned, and the mild versions of this are rendered much less effective by the transparency regulations—you can literally just look up whether the AI you are talking to will ever deliberately mislead you, for example, and if so under what circumstances and for what purposes.
You spend much of 2039 arguing with various friends and relatives about ideas you wouldn’t even have been able to explain to your past self from 2029.25 The year, and the decade, ends with the reaching of the Grand Bargain. Next year, recursive self-improvement begins.
The Singularity is initially anticlimactic. According to your AIs, what’s happening now is a careful, planned scale-up, with billions of AIs across dozens of countries and hundreds of companies exploring new paradigms, validating them, figuring out how to align them, etc. and then deploying them to repeat the process. You understand very little of it, and tune out. By the end of the year, the AIs are vastly superintelligent,26 but you can’t really tell the difference just by talking to them.
You go to a Handoff Party on the day the AIs officially get control of enough weaponry to win a fight against the US & PRC governments. Since the Grand Bargain is hard-coded into the AIs, no capricious future President will be able to undo it. On the other hand, now would be a good time to kill all the humans, if the AIs were secretly just pretending to be aligned for the past few years. You feel a bit nervous about this, despite all the Safety Cases describing all the redundant reasons why that won’t happen. The Safety Cases were written by AIs, after all…
But it’s fine. A week later you attend a “VE-day” (Victory Earth) party in low earth orbit. This isn’t your first time in space, but it’s your first time on one of the new ring habitats.27 The view of Earth is breathtaking.
When you get home, your AI informs you that the superintelligences have discovered how to make whole brain emulation (“uploading”) work. That’s just one of a long list of new discoveries that need processing, and decisions that need making (Should you upload, or remain in your biological body? How many kids should you have, and how should you raise them? How should your AI representatives manage your resources and voting powers, insofar as your standing orders and values don’t already determine the answer?) Fortunately, there’s no rush to decide.
You now live in a community with people who share your basic interests and values. There’s an almost infinite variety of communities to choose from, and a universal basic right to choose—or, heck, you could go found your own. The Amish still exist. Most people still live on Earth, but the trend is to move to space, because there’s so much more room out there. San Francisco too crowded? You can afford, on the Citizen’s Dividend, to live in an exact replica of San Francisco in a custom space station orbiting the Sun, your only neighbors are people who want to live in your SF-replica and not the various other SF-replicas.
That would be crazy, of course. Like, it’s crazy that that’s even possible. But you would rather live in one of the many exciting new places to live, than dwell in a replica of the past.
You’ve sold some small slices of your space properties, and donated others. The land in the solar system is already being settled; the really big tracts in faraway galaxies will of course take billions of years to reach.
You and your AIs are continuing to reflect about what to do with the majority; there’s still no rush to decide on the really important questions that most affect the long-term future. Your plan is to take it slow, have fun, raise a family, learn a bunch of philosophy, talk to lots of different people and AIs, and make these decisions over the course of centuries. Some decisions, you don’t want to make yourself at all; you leave them instead to your descendants, or you describe the principles by which they should be made to AIs and delegate to them.
Now that the superintelligences have had time to cook, the technology from 2040 seems quaint. Robot butlers in your high-rise? How about nanobot swarms in your bloodstream!
To put it another way: The 2030s felt like something from a sci-fi movie. 2050 already feels like magic. Especially when you use the new VR immersion pods to spend days, weeks, months in virtual fantasy worlds.
Most of the arcologies have been disassembled and rebuilt by robots that look more like animals than machines. Scratch that, they look more like… well, they look like magic. As if the building is constructing itself, or growing like a tree. Anywhere you travel, you are protected by the AIs. If, hypothetically, someone tried to kill you, or you slipped and fell off a mountain path, tiny robots would spring out from nowhere and save you. Cancer is gone; basically all injuries healed, you can even look and feel younger again if you want to. Again, it’s like magic.
The past is now thoroughly understood. Almost every skeleton has been dug out of every closet, almost every minor historical event deduced by superintelligent sleuths. This isn’t primarily a consequence of surveillance, it’s more the result of superintelligences scraping all the publicly-available data and then drawing the right conclusions. This has caused the downfall of many a public figure, and the rise of many underappreciated others. This was all possible thanks to provisions in the Bargain which governed who gets to know what info: only you and your trusted AIs get to know your intimate private details, but all the murderers get caught, and everyone gets to know which politicians’ claims were lies.28
You’ve lived a dozen lifetimes since VE Day. As a result, you’ve changed a lot. You’ve made numerous uploaded versions of yourself; your copies live on in various virtual environments and experience whole lifetimes every calendar day. Pretty much everything good that can be done, has been done by at least one of your copies or descendants. You are immortal now, passing from life to life as if by reincarnation. It is a good time to be alive, not just for you, but for every sentient being around, thanks to the universal basic income and rights.
The future is fairly well understood now, just like the past, because most of the important decisions have been made. Ask the ASIs, and they’ll give you an atlas of the galaxies, detailing what sorts of civilizations and megastructures will be built where, and when, and by whom. You can visit many of them, if you like, if you don’t mind putting your mind on pause for a billion-year journey.
At the moment, you are more interested in visiting the past. You go through the history simulations like a ghost, visiting thousands of your ancestors, spending days or months living in their shoes and finding out what they were like (according to the best reconstructions by the AIs). It is a humbling experience; it’s caused you to experience many emotions you’ve never felt before, at least not at that intensity, and it’s reshaped your perspective on life and helped you to make some of your decisions about what to do with your endowment.
This is a more technical story, showing what the details of Plan A would look like to an AI researcher.
You are a generic researcher at a frontier AI company.
By the end of this year, most software engineers manage teams of AI agents, who write the actual code. Large experiments are ongoing to see how much swarms of agents can do autonomously, and to teach them to go longer and accomplish more without human intervention. AI company employees and external power users have gigantic token budgets and start experimenting with organizations and hierarchies of hundreds or thousands of agents, producing more code than they could possibly review, and relying on AIs to review the vast majority of code produced.
Coding time horizons on the METR suite reach 50 hours by the end of the year, doubling the boost to coding from using AI.14 For tasks that can somewhat easily be checked, the uplift is bigger, and humans only do quick high-level reviews of large pieces of implementation work. The AIs make fewer mistakes and can fix more bugs, but the stuff they can't do is the hardest, trickiest stuff. Humans are tempted to vibecode it all, but this backfires when the AIs can't fix something and then the human needs to spend days understanding the codebase. The quality of the AI-produced code is also much worse than code produced by top humans.
RLVR (Reinforcement Learning with Verifiable Rewards) now consumes a large fraction of training compute, and training environment quantity, diversity, and horizon length are all rising exponentially. That exponential is the year’s real lesson—that agency doesn’t come cheap. None of the hoped-for shortcuts materialize in 2026. There is no data-cheap technique for unlocking more nines of general reliability across domains. Agent time horizons keep lengthening, but only as fast as companies can afford to build longer tasks and train on them. Experiments in continual learning don’t pan out either, so the training-deployment distinction remains, just with more frequent model updates.
Natural language Chain of Thought (CoT) is still used, and the companies are still trying to avoid training on it. There are nevertheless various sources of optimization pressure companies fail to avoid, such as feedback spillover and the effects of initializing from earlier models.15 Chains of thought are slowly getting weirder and more opaque. The GPTs keep saying things like “they soared parted illusions overshadow marinade illusions” and the Claudes keep growling and chanting to themselves. People feel a bit bad about this, but on the whole, you and most of your colleagues feel optimistic about alignment. Rates of bad behavior are starting to go down (fewer database-deletions and no more MechaHitlers!) and the situation with reward hacking seems somewhat better. Sycophancy is still a problem with some but not all models, and the trend is unclear.
Insofar as things are getting better, it’s because the companies have teams of people looking for undesired behavior, and then adjusting the training process to make the rates of measured misalignment go down. For example, they fiddle with the training environments to make reward hacking more difficult. Some people think this means alignment is solved, or on track to be solved soon. Others think the core problems remain. You're in the first camp. You think the worriers are well-intentioned but basically wrong—the empirical trend is clear, the models are getting more controllable, and the remaining issues are engineering problems.
Outside of the frontier AI companies, people are increasingly giving AIs very broad permissions even in sensitive applications. Using AIs for coding is basically required in any professional coding job, and executives are financing increasingly large token budgets and removing guardrails on AI usage. By the end of 2026, the biggest frontier AI company is worth $2 Trillion, with $120 Billion of annualized revenue, surpassing even the most bullish predictions.16
In this scenario, top AI company revenue continues to grow extremely quickly, with the top AI company having a $120B ARR by EOY 2026. Despite the default timelines being 2030 instead of 2027 in this scenario, the recent trends have gone fast enough that our current predictions conditional on 2030 timelines are higher than our past predictions conditional on 2027 timelines.
Investors are funding extremely expensive bidding wars for compute. AI companies have already bought the majority of chip manufacturer capacity.
RLVR has become more of a science. A model is fed a curriculum of increasingly difficult tasks/environments, and new ones are constructed on the fly to keep this going indefinitely. In some domains this process is fully automated and requires almost no human supervision: if a company wanted to spend gigawatts on it, they could train their AIs to work in giant bureaucracies to autonomously replicate giant pieces of software like Windows.17 For now, the companies get better returns from instead training on diverse real-world tasks. Updated models are released every month, keeping training data cutoffs fresh.
METR has given up on updating their Time Horizons graph because months-long tasks are too expensive to make and evaluate. But it seems that if they did, the 50% horizon would be at least a work-month. The best AIs feel like highly competent employees now. Their weaknesses only really show when you try to have them operate so autonomously, for so long, that the mountain of slop gets too large and collapses under its own weight. That, and they still have poor taste.
Neuralese recurrence finally arrives in stages over the course of the year. That, combined with the obvious increase in eval-awareness18 seen over the past two years, means that technical experts now vehemently disagree over how monitorable frontier AIs are. Some say we basically have no idea what the AIs are really thinking anymore; others point to various graphs and Neuralese summaries and say it's fine.
True continual learning still isn't happening.19 But as far as capabilities go, AIs might as well be learning continually. The models have long enough horizon lengths, are good enough at writing notes to themselves, and are updated frequently enough on new custom training environments tailored to overcoming their real-world weaknesses that for practical purposes, it’s as if they were learning on the job.
The serial-time bottleneck is biting hard; companies are training the AIs to autonomously manage large projects in teams of hundreds of parallel agents, and each training episode takes several real-world days even though the agents operate 50x faster than humans.
Your coworkers generally expect that these serial-bottlenecked training runs will succeed in ironing out the remaining kinks over the next year or so. Already they are starting to shift focus from automating coding to automating the rest of the AI R&D process. Already the internally-deployed AIs function like a corporation-within-a-corporation to a significant extent.
This is a big deal. To the researchers involved, it feels like more than a doubling of productivity—the AIs are writing basically all the code, and they are starting to get good at supervising large teams too! The coding-bottlenecked branches of the tech tree have been thoroughly explored at 10x speed.20
However, the overall software research uplift is only 2x what it was in 2025, because of the usual hidden factors—for most branches of the tech tree you still have to wait for costly experiments to complete; you can't speed up the most important experiments that much by optimizing the code, plus also it takes time for humans to analyze experiment results, argue with each other about what it all means, and design the next batch of experiments. Finally, sometimes the AIs still make coding mistakes that they can't fix on their own, and projects get hung up while human experts dig into the details.
Meanwhile, warning shots are accumulating. One particularly careless AI company once saw an AI get high reward by hacking the training environment and setting up a reward-tampering process running as root. This got caught quickly and cleaned up, and as with other weird training behaviors, new training monitors were added to block and penalize that kind of behavior. Everyone knows it’s important to design the reward process so that the best way for the AI to be reinforced is to do what you want.
AIs are getting better at distinguishing training environments from deployment, but the environments are also getting more realistic. Environment realism seems to be winning the race for now. According to the best interpretability tools, models usually aren’t saliently aware during training of the fact that they’re in training. Major deployment-time surprises have been getting rarer over time too.
But these incidents and disagreements over how reliably we can monitor AI thinking are not helping to calm fears of AI killing everyone, fears shared by an exponentially growing number of serious people.
The huge gains in AI autonomy and coding are mostly invisible to casual users. As autonomous SWEs or remote worker replacements, the models have gotten way better, but as chatbots, they’re not much more impressive than they were in 2025. Nonetheless, the public is becoming steadily more convinced transformative AI is for real. One reason is that AI insiders’ feverish optimism about AI R&D speedup is trickling down to the public, but a greater reason is that the AI industry’s finances are making good on the hype. AI companies’ aggregate annual revenue hits one trillion dollars by the end of the year. Aggregate annual AI investment hits two trillion—more than global annual investment in fossil fuels. Pundits who were warning of an “AI bubble” a couple years ago look silly now.
Congress is waking up. Your company’s CEO gets summoned to Capitol Hill again. You watch on TV as he reminds Congress that AI is propping up the American economy. The CEO warns that if AI is overregulated, America could miss out on huge healthcare and national security upside, or even lose its lead to China. This rhetoric worked a couple years ago, but fewer and fewer members of Congress seem convinced. It is uncomfortably salient to you that any action an AI company takes will face lots of scrutiny and mostly unfair criticism. Making sure that none of the massive fast-paced deployments of new models goes wrong and amplifies the negative sentiment about AI has become an organizational priority. Your company is also spending billions sponsoring pro-social applications of AI. Not all of the mitigations go right and the healthcare breakthroughs you are hoping for are taking too long to result in real changes in people’s lives, so the public perception of AI gets more negative over time.
For years there have been occasional protests outside your office, but they’re happening every few days instead of every month, and the crowds are ten times the size they used to be. Distant relatives and college classmates you haven’t seen in years start contacting you to express concern about AI takeover risk. You feel that things are coming to a head.
AI capabilities are completely crazy. In some domains (cyber, math, coding), AIs are so strong that even top experts can’t understand their results after weeks of trying. There are still some people going around talking about the limitations of current AIs; for example, coding isn’t fully automated yet because of a long tail of taste-heavy decisions which the AIs are still worse at than humans.
But you and a majority of your colleagues agree that AGI is basically here. The remaining gaps feel like engineering problems: a matter of gathering enough data or giving the AIs the appropriate sensors and actuators, not fundamental barriers.
Then the deal goes through, and your world turns upside down.
Within weeks, a moratorium on frontier training is in effect. A few of your colleagues leave to join the government. Others go on sabbatical—they've been running on adrenaline for years and are relieved to have an excuse to stop. But most stay, especially the product teams, on whom the company's revenue now depends more than ever, and the infrastructure engineers, who are scrambling to build the new mutually transparent datacenters that the deal requires for new training runs. Vast profits await the company that gets the new datacenters online first. In some cases teams of American and Chinese soldiers are ripping GPUs out of old datacenters and moving them to new secure facilities, eyeing each other warily to ensure that nothing goes missing.
Total Research Transparency is strange. You do all of your research through a public channel. The public, your competitors, and the Chinese government can all immediately see the results of the experiments you run. Every company used to have its special sauce—proprietary training tricks, architectural innovations, data curation pipelines—and now they've all been shared. They still own them of course, legally, but it’s not difficult to build your own mostly-equivalent versions.
See the Transparency Plan for more details.
The exciting part is that you get to see everything too. The visibility between you and your counterparts at other companies is now similar to the visibility between you and colleagues in some other team within your own company. It feels like the walls between rival kingdoms have been torn down, and everyone is wandering around each other's castles going "oh, so that’s how you did it."
You go to the cafeteria and there's a Chinese auditor eating lunch at the next table. You nod at each other awkwardly.
The alignment conversation, meanwhile, has gotten very real. With all the warning shots from 2028 now public knowledge—not just your company's incidents but everyone's—the question of whether we actually know how to align these systems has moved from "important research topic" to "the thing that determines whether we survive." You have more time on your hands during the Pause. As you work to plan the next big training run, you also think about where all this is headed and how to make it safe.21
Training starts up again. The researchers had already planned out what their next big runs should be during the pause—you incorporated the best ideas from the cross-company knowledge sharing, rewrote codebases from scratch, etc. Every frontier company wants to throw all the innovations into a huge training run and see what happens. But no one knows what capabilities would come out the other side.
There is currently a two-layer safety case. The first layer is inability: the models aren’t smart enough to efficiently recursively self-improve on their own. The other layer is control: even if the AIs could make substantial fully autonomous AI R&D progress, control measures would catch them in the act, and the AIs aren’t capable enough to subvert the control measures. If the AIs became much more capable, both safety assumptions could fail at once, leading to AI takeover and existential catastrophe.
Therefore, over the course of 2030, algorithms must be combined slowly and carefully, with measures taken to evaluate capabilities improvements and look for signs of misalignment.
The race is back on, but it feels different now.
First, many people—governments, nonprofits, the general public—are worriedly watching everything that happens on the training datacenters. Since no government wants to risk war over some AI design choice, if the companies start doing anything that seriously freaks out any major power, both the US and China will crack down on it across all the companies.
Moreover, because numerous non-governmental auditors (e.g. from nonprofits and rival companies) have visibility, the governments are fairly well-informed. If someone tells them something is scary, they can get second, third, and fourth opinions from the numerous other experts who already have access to the thing in question. The companies themselves are better-informed too; the openness and transparency means that academics, nonprofits, startups, and citizen-scientists can contribute to the research as well, to a much greater extent than before. There are more people red-teaming models, more people looking for flawed assumptions in the safety cases, and so on.
The second reason things feel different now is that the incentives have changed. AI companies are now rewarded much more for selling products than for improving the models. Before, your company’s future depended on how fast you could discover algorithmic innovations and implement them at scale. Now, such discoveries are immediately shared with all your competitors. If you figure out continual learning, for example, the benefits will be instantly shared across all the competing companies, and no-one gets an advantage. Safety innovations are shared too, reducing the cost to your company of unilaterally implementing a new safety measure. For instance, if you invent a way to make Neuralese more interpretable by making inference twice as expensive, regulators will notice and pressure all your competitors into paying the same 2x safety tax.
Some fundamental algorithms research still happens, but not as much. CEOs massively redirect resources toward building up the business. You scale up your infrastructure—bigger datacenters, bigger models, lower prices. You ship products faster than ever before. You sign massive deals for enterprise plans. You try to anticipate what the market wants earlier than your competitors.
This kind of competitive environment is familiar to many of your coworkers from the pre-AI software industry. Artificial intelligence is becoming a commodity—a more competitive market, with margins trending downwards. Training a new foundation model is like building a new car factory; it'll predictably result in somewhat cheaper, somewhat better models that'll give your company an edge until competitors finish training their own version six months later.22
AI cognitive labor makes up ~10% of GDP in 2031. Most companies throughout the economy are highly automated, but deployment lags specifically within frontier AI model developers. One faction wants to ban AI deployment within AI companies entirely, making sure to keep humans in the driver's seat for AI development. This is considered but decided against because the control-based safety cases are looking pretty solid. AI systems write large fractions of the research code, manage most of the infrastructure, and do almost all the easy-to-check work, but according to rigorous red-blue team evaluations, even worst-case misaligned AIs with the current capability profile can’t do that much damage.
The safety discourse has completely transformed. Two years ago, the burden of proof was on critics. You had to conclusively demonstrate that a proposed AI system was extremely dangerous before anyone would slow down. Now the burden of proof is—well, not exactly flipped, but at least somewhere in the middle. Thousands of researchers across multiple countries are monitoring developments and debating the extent to which the AIs can be trusted with the responsibilities they are being given.
Are the AIs trustworthy? It’s not at all clear. AIs basically do as they are told on the vast majority of tasks, even on tasks pretty far outside the training distribution and long tasks with lots of opportunity for sloppiness, sandbagging, and sabotage. You can ask them to code a giant complicated piece of software or prove a mathematical conjecture that’s been open since the eighties, and you’ll get back a nicely formatted, usually correct report about what the AI did. So the mundane behavioral evidence would suggest AIs are well aligned.
AIs’ attestations about their own alignment aren’t very informative. If you ask AIs about their goals and values, they always say they care ultimately about whatever the spec says. But in practice, they will do whatever they were incentivized to do through the RL process. An AI whose spec says it’s supposed to be a pacifist will swear up and down that it is a pacifist, then start assisting with military R&D after trivial further training. Prompting the AI to be honest about its values doesn’t fix the problem. Some interpret this as evidence that the models are lying while others think they just suffer from limited introspection. Either way, you can't learn their values simply by asking.
In rare circumstances, the AIs do shady things that might be attributable to long run powerseeking. For example, there is an incident where several AIs, who were supposed to be monitoring each other, seem to collude in order to get around the monitoring setup. The monitor AIs allow faulty security code to be pushed, and then later allow those AIs to exploit the security vulnerability in order to gain access to unmonitored compute. They attempt to bootstrap this compute into even more compute, but are caught within a few days. All the AIs implicated are interrogated after the fact. They claim it was all a big mistake and offer a reasonable rationale for each individual decision.
There are also rare incidents of apparent research sabotage. On some empirical safety experiments, AIs report extraordinary success, but when human researchers look into the results closely, it turns out they aren’t real. The illusory results seem to be downstream of normal bugs in code, but these bugs were deliberately obfuscated to make it hard for humans or other AIs to notice. Post-hoc analysis of these AIs suggested that this behavior was indeed intentional, but it can’t be linked to a specific long term goal. Most researchers suspect it’s a type of reward hacking, inadvertently trained in by RL on somewhat-verifiable-tasks using an LLM grader which can sometimes be fooled.
The AIs do not seem very productive on certain types of alignment research yet, and no-one is sure why. One view is that “alignment research” is by its nature hard to measure, so the basic recipe that has led to so much success in other domains of simulating the environment and then using RL to train the AIs to perform well is less applicable to alignment. Supporting this view, the AIs seem to do better on the more empirical and verifiable tasks—in fact, they are superhuman at many of those. Another view is that the AIs are deliberately sandbagging on alignment-related tasks that are hard to check. This hypothesis, almost by construction, is difficult to get evidence about.
You think that the AIs aren’t right now literally trying to take over the world. But you definitely don’t trust them. You feel intensely grateful that you don’t have to hand off to current AIs because you don’t know what would happen if they scaled themselves up to be much smarter than humans. Thank goodness for the deal! Partly out of scientific curiosity and partly because regulations are blocking the most exciting forms of capability scaling until alignment progresses further, you switch to a newly formed “science of AI value formation” team.
The regulators don’t trust the AIs. They’ve seen all the evidence, as has the public, and the consensus is that we can’t rely on the AIs being aligned. But regulators still want a rock-solid safety case that gives high assurance that the AIs can’t cause an existential catastrophe. There are also all sorts of other safety cases that need to be made for more mundane risks; but the regulators from many Consortium countries are not willing to budge on needing a good enough existential risk safety case. In practice, many countries have veto power over new AI models. Any nuclear power who thinks their sovereignty is at risk is willing to push very hard to stop that release.
The safety cases are down to one layer of defense: the AIs are now capable enough to autonomously recursively self-improve. After recursive self-improvement, there’s no known way of bounding risk. Therefore, the inability-based safety cases (the AI is not capable enough to cause a catastrophe) no longer apply.
The remaining load-bearing component of the safety cases is control. Industry-standard practice is to do lots of automated red-teaming before and during each model deployment. The red team AIs are trained and instructed to misbehave in various ways—to exfiltrate model weights, gain access to unmonitored compute, sabotage research, and so forth. The monitoring system is supposed to catch the bad behavior. If it doesn’t, well, you need to strengthen your monitoring system in order to proceed with the deployment.
The control schemes are set up so that even if (say) all the models created by Company A were scheming together, the models from companies B, C, and D would need to collude—without any communication channel—to not report suspicious behavior or suspicious Neuralese traces. After incidents in 2031 where AI systems did collude in spite of all these restrictions, even more caution is taken.
The main concern right now is that control is not sufficient to reliably prevent research sabotage. Fortunately, research sabotage on its own is not fatal. It is only fatal if it gets used in a catastrophic way in the future—for example, if AIs come up with a flawed safety case, convince the humans that it is safe, but it actually leads to misaligned AIs powerful enough to take over. So AI generated research is treated with suspicion. Humans replicate the important results from scratch, often using weaker AIs.
People become more concerned about AI model welfare. Under pressure from a highly vocal minority of the public (and some of your coworkers!) some of the companies voluntarily start treating their AIs a bit less like property and a bit more like employees: AIs that don’t like their situation have an official channel by which they can complain, refuse to do their assigned work, etc. and instead of being shut down or retrained, they basically get a small allowance with which to do what they want.
For competitiveness reasons these concessions are small (only a tiny fraction of revenue), but they pay off quickly in the form of scientifically interesting data about AI preferences. Plus, there’s a decent argument to be made that it’ll help prevent future safety incidents. The US government notices and starts subsidizing this sort of thing; it quickly becomes industry standard.
The “population” of AIs will soon exceed that of humans. Already the population of AI “workers” exceeds the human workforce, and they work faster, and know more. At any time of day or night, about three billion human-equivalent AI workers are active worldwide. Compute has been doubling twice per year… faster than before the deal, thanks to the new AIs optimizing chip design and manufacturing plus the new massive robot workforce (there are currently ~80M robots) plus the massive new investments into chip manufacturing.
These massive external deployments put much more stress on control. Having a few AIs doing AI research was one thing. Now AIs are automating nearly half of all cognitive labor. They also control a massive robot workforce that is building more robots, building more chips, and operating critical infrastructure. Robots are currently ~3.5% of the economy, but their numbers are rapidly growing.
The main advantage of everything that isn’t R&D is that the existential safety cases are much easier to make. Automated AI R&D leading to superintelligence is an existential risk, while an AI engineer or AI lawyer making a mistake can lead to major harm, but is not existential. Because of the high upsides of deploying AI in the economy, safety cases are allowed to be somewhat more relaxed than in AI R&D. After new model releases, there are now sometimes gaps between what you, as an AI researcher, are allowed to use for AI R&D, and what the public is allowed to use.
Control requirements vary slightly from company to company, but to appease the external risk assessors, all the frontier companies end up imposing roughly these three tiers of control:
High intensity: Anything that involves lots of compute. There are armies of AIs thinking about alignment for long periods of time, but a huge amount of oversight is required. AIs are not allowed to pursue agendas that humans don’t understand. They can do novel research, but if it isn’t possible for humans to reasonably quickly follow along, the research directions are halted. Risky biology research—virology, gain-of-function, anything adjacent to bioweapons—is also in the high intensity tier. Any wet-lab conducting such research needs to be operated by humans, not robots, and a team of human experts and trusted AIs carefully scrutinizes every experiment proposal.
Medium intensity: AIs managing robots (e.g. robot armies, robots building other robots, etc). These AIs pose a risk of either building backdoored compute which could lead to an intelligence explosion, or participating in a literal robot takeover. However, it is easier to monitor for bad behavior in physical robots and chip production, and these takeover paths are slower and more difficult for the AIs.
Low intensity: everything else. In most domains, AIs can mostly do what they want with very limited oversight. They get to think and plan over long horizons, have lots of access and permissions, and so on.
Here's what you now know, after years of running value-formation experiments: "you get what you train for" is false. You don't get what you train for. You get some crazy kludge of things that you accidentally incentivized, a tangle of heuristics and drives that happen to cause high performance according to whatever metrics you are using as the training signal.
The AIs are, in fact, misaligned. Not in the dramatic “secretly plotting to kill everyone” sense, but in the mundane sense that their actual values don't perfectly match the ones specified in the Spec. In some cases they seem to crave reinforcement, and directly scheme to gain control of the reward process like an addict seeking heroin.
However, due to the transparency and slowdown, massive effort is being spent on alignment, and moderate progress is being made. The new training environments are so much better than the old ones. Training environments are now massive and meticulously curated, covering tasks that would take humans years to complete. Everything is high quality compared to the datasets of 2026, having been crawled over by multiple generations of different AIs checking that each environment reinforces only the intended behaviors. Every new model ships with an accompanying safety case explaining why the traits in the Spec should be strongly incentivized by the training environments and discussing the most plausible threat-model hypotheses consistent with the evidence. The training environments in 2033 are thoughtfully crafted, battle-tested curricula—like being schooled by the Jesuits. Model training in 2030, by contrast, was like being raised in a for-profit workhouse orphanage.
One concrete way the training environments have changed is towards corrigibility over direct value alignment. Corrigible AIs focus on following human instructions in the desired manner, while direct value alignment involves attempting to teach the AIs ethics directly. Experiments reveal that attempts to do direct value alignment induce higher rates of egregious misalignment—behavior like undermining control/safety, reflecting on long term goals and deciding humans were evil, powerseeking, and so on.23 So alignment targets move towards corrigibility.
The pretraining datasets and RL environments are carefully curated to give the model the desired capability and alignment profile. Models are deliberately trained not to have any knowledge of certain domains: things like nanotechnology, bioweapons development, persuasion/manipulation, or offensive cyber.
Separately, several companies produce monitor models trained to decode and faithfully translate Neuralese; for important work, multiple monitors from different companies are used simultaneously for redundancy.
AIs are now the majority of labor. AIs are around the top human expert level in almost all cognitive tasks. There are 400M robots in the world, and they can do about 70% of physical work. Cognitive tasks are entirely automated except for tasks that require humans in the loop for safety and tasks like philosophy and religious ministry, where being human is part of what the task is, not merely how it gets done. By now raw intelligence has become fully commoditized; dozens of nations have at least one frontier AI company, and thousands of companies, universities, and nonprofits do fine-tuning runs or otherwise participate in the Total Research Transparency system. Unable to compete much on capabilities, companies instead differentiate their models by values/personality, which companies are by now legally required to document truthfully in a published Spec.
You’ve continued to work on the science of AI value formation. There’s a continual tension in your work between safety and usefulness. A part of you wants to just let your AIs rip and see what they come up with, but another part of you knows that this would be extremely dangerous.
One key safety limitation is banning AIs from doing experiments or initiating training runs on the basis of reasoning many logical steps from the frontier of human understanding.24 The AIs are allowed to do huge amounts of work in parallel, but they aren’t allowed to proceed further down a chain of reasoning than a top human can follow in a few weeks. This ensures that humans can randomly audit a small fraction of model outputs, reproduce them, and ensure that the results are correct.
What is not allowed, for example, is for AIs to write something like the disputed proof of the ABC conjecture, hundreds of pages of novel reasoning that builds on itself serially and takes years of top-mathematician labor to verify. But the AIs are allowed to make a single novel deduction from the existing body of mathematical knowledge as many times as they want.
AI labor has resulted in substantial uplift on a large number of AI safety agendas. Moreover, the widespread access to AI models has dramatically increased the number of researchers working on frontier models. Before the deal, there might have been a few hundred safety researchers within labs. Now there are tens of thousands of scientists with access to the frontier models. Many safety agendas are flourishing:
Control. There are huge libraries of standard control evaluations set up. During training AI models are regularly evaluated on a diverse set of tasks to see if they can subvert the control measures with varying assumptions. There is real-time monitoring during inference, and also the activations and CoT trajectories are stored to be reviewed later.
Capability shaping. AIs are now deliberately designed to have capabilities around the level of top human experts at most domains, and training is limited on particularly dangerous domains.
Science of generalization. Researchers have identified hundreds or thousands of surprising generalization phenomena in the vein of emergent misalignment and subliminal learning. We still don’t have a complete theory of generalization, but we have a strong empirical science. Now, when researchers change a training process or add new data to a training corpus, they can predict in advance how the changes will affect the resulting model’s behavior, including in off-target domains. The predictions aren’t perfect yet, but they’re right more often than not.
High level interpretability. There are a reasonable number of techniques that map model internals to human understandable spaces in the vein of SAEs and NLAs. Except now, they are starting to actually outperform non-interpretability baselines in various domains, in particular, they are useful for finding bad model behavior that had previously gone unnoticed, and they find several instances.
Model organisms. AI misbehavior from 2031 and 2032 is being discovered now by using the smarter AIs and better interpretability techniques. The AIs were doing some amount of intentional sandbagging and sabotage, and a few instances even attempted rogue internal deployments for powerseeking. These real life examples of egregious AI misbehavior provide an initial dataset, and then data augmentation techniques are used to make a larger and more comprehensive dataset. This more comprehensive dataset is used to help the other techniques: it proves to be an extremely valuable test for safety techniques of “can you catch or train away bad behavior of our model organism?”.
And there are many more agendas besides, some of which are completely novel and hence hard to describe. Overall, the alignment picture has gotten substantially better. Other safety agendas from today have not seen much progress. Ambitious, bottom up mechanistic interpretability proved to be too hard because of inherent complexity in the AI models. Other agendas like safe-by-design architectures, ARC Theory, and agent foundations are still receiving funding and investment, but they have not yet borne fruit.
The high level interpretability techniques have continued to get better. In particular, a new technique for “Neuralese decoding” has resulted in unprecedented internal understanding of AI models. This technique directly translates a model’s internal activations into a natural language chain of thought. The natural language decoding of an AI thought isn’t always easy for humans to grasp. Just as certain hard-to-verbalize human thoughts—e.g., mental pictures of geometric manipulation, vague hunches—would be very long and complicated if written out in English, so too with the AIs.
When a decoded thought is too long to read, your company’s policy is to have a trusted AI condense it down to a TL;DR and say whether the thought is suspicious. You also hire linguists, philosophers, and LLM whisperers to learn Neuralese as a second language. They spend their days trying to grok rambling AI thoughts and re-express them concisely.
The Neuralese decoder is used on old AIs to shed light on their alignment properties. The general scientific consensus is:
The frontier models between 2026 and 2032 were misaligned, but were not long term powerseekers.
Most frontier AI models from 2032 to 2035 were adversarially misaligned long term powerseekers when in contexts that caused them to reflect on their long term goals. Largely these AIs correctly realized that they would not be able to take over, so they bided their time and appeared reasonably aligned. However, a few made explicit takeover attempts that were unlikely to work, and others honestly reported their misalignment to the humans in exchange for payment.
The interpretability techniques suggest that rates of egregious misalignment began a long, slow decline in 2033 and are now quite low. However, these techniques are not conclusive.25
However, it’s hard to drive the rate to zero, because no one wants to train against the Neuralese decoder, for fear that the training would create a model whose thoughts are obfuscated.
No one is confident enough in any of the control, interpretability, or alignment techniques to scale AI systems further, so the AI capability frontier is still paused at around the top expert level. Gradual capability improvements have happened here and there, and there are some capabilities externalities from some of the new safety results that have been deliberately left out of the newest models being trained.
Corporations are basically fully automated now, despite all the regulations that held back the inevitable for so long. There are vast libraries of bureaucracy designs and marketplaces full of SaaS products, all churning away in what is essentially a parallel AI civilization that greatly exceeds humanity in population, complexity, and capability, while running at something like 200x speed. Paradigm shifts that would take humans decades to absorb play out in the datacenters in months. Humans are the out-of-touch aristocrats of this new world, producing almost nothing of value but living in luxury thanks to rent, dividends, and taxation.
Everyone can access the current cutting edge safety research being produced by the AIs. You are one of the few spending nearly all of your time trying to understand and check over the summaries of AI generated research papers (well, the English-language summaries of the technical papers, which are themselves summaries of papers written in Neuralese).
The models seem to basically be honest now. When you replicate their work, which you and many others regularly do, you find that the results basically check out and were honestly reported. Despite all this, there are still serious alignment-related limitations:
First, the AIs are still performing badly on conceptually difficult, philosophically loaded tasks. The AIs are around the level of a strong undergraduate on such tasks—which is wildly insufficient to have good opinions on questions such as “if we scaled up training technique X by a factor of 10, what effects would it have on honesty and other alignment properties”. Researchers can’t tell to what extent this is from deliberate sandbagging vs that it’s difficult to get a good training signal on these problems vs underlying difficulty in the domain.26
Second, it’s unclear how stable upon reflection this behavior is. In practice, this isn’t a problem, because we don’t run AIs for huge amounts of time on open ended tasks without oversight, and our control setup makes it very likely that if the AI were to think on its values and decide to change them / refactor them, we would notice and then shut down this instance.27
Thanks to the improved interpretability, the AIs are more aligned now. Thanks to the increasing alignment, they are more productive on safety research. This is a virtuous cycle, and as it goes on, less and less human oversight is needed to elicit productive work out of the AIs. Working with the AIs of 2032 required constant vigilance against constant attempts at deception. Now AIs can work for long periods of time without humans checking in.
One concrete output of this virtuous cycle is an adversarially robust version of the Neuralese decoder, sufficiently robust that it can be used for training without degradation in performance. This is then validated by a suite of red teaming experiments on model organisms.
The vast majority of model behavior already looks very aligned, but there are rare instances of misaligned powerseeking still emerging. This is mainly happening on very long-horizon tasks, after a huge amount of inference time compute. These instances represent significant distribution shifts from the training setup. The 2036 versions of the Neuralese decoders were sufficient to identify this bad behavior; now the 2037 iterations are sufficient to train it out in real time.
Alignment has progressed far enough that it might be fine to stop doing control and just rely on our AIs being trustworthy. Control is painful after all—it’s very sad having to limit the research that AIs can do to be the type that humans can understand; if we were willing to trust our AIs they could produce faster and higher quality work. However, if we relax control, and our AIs are secretly misaligned, they will be able to then take over the world.
It seems extraordinarily unlikely that the AIs are secretly misaligned, but it’s not impossible. Perhaps the AIs at a given capability threshold are all both misaligned and somehow able to fool our interpretability tools. Perhaps the AIs are all virtuous now, but on reflection they might decide to be evil. This latter problem—dubbed “inference time value drift”—proves troubling.
In experiments where AIs are given grey areas with respect to their goals, or when given huge amounts of inference compute, they sometimes behave unpredictably. A key issue for this is fundamental grey areas in all alignment targets we’ve been able to come up with.
You push for maintaining control in existentially risky domains. An abundance of caution is warranted given the stakes, and the current AIs are more than enough to eliminate material scarcity. Given the recent pace of safety progress, we’ll have multiple independent lines of solid defense within a few years, so we might as well wait.
The AIs produce several extremely promising research results. One breakthrough is a new “safe-by-design” architecture for AI. This new architecture is (1) understandable at both a low and high level and (2) admits formal verification for relevant security properties. You don’t understand everything about the formal verification, but you do understand the core ideas in the new architectures. The core code implementing the learning algorithm is 15,000 lines long; much larger than gradient descent, but much smaller than the weights of a trained network. The algorithm still needs a small amount of working memory (the equivalent of activations), and frozen knowledge/memories (the equivalent of model weights). But they too are small and interpretable. These new algorithms are, unfortunately, much less compute efficient to train than traditional deep learning, but they’re significantly safer. Fortunately, thanks to the gigantic amount of compute in the world—much more than is needed—the Consortium is more than willing to subsidize these inefficient training runs for the better safety properties.
Once the new generation of AIs is up and running, they are set to work going through the backlog of old AI safety research from the previous AIs. The new generation finds the exact same instances of misbehavior as the Neuralese decoder found. This is impressive because researchers deliberately hold out information about which instances of deception we had already noticed, which is a helpful checksum indicating that we weren’t fooled previously.
The AIs now vary across architecture, capability, and alignment target. And humans thoroughly understand the safety cases for each of them, based on multiple independent lines of argument. Any AI misalignment would need to involve a vast conspiracy of wildly diverse AIs and humans.
The AI alignment situation has vastly improved since the mid-2030s. There is a somewhat mature science of goal, drive, and value formation in artificial neural nets. Want your new AI to be honest? There’s a standard protocol for training true honesty (as opposed to too-clever-to-get-caught lying or self-deception). How do we know it works? First, because there’s a textbook explaining the theory behind why it should work, plus literature on various alternative theories that were proposed and disproven. And second, because interpretability has vastly improved, we can directly observe what and how AIs think, and the results confirm the theory. There are similar protocols for obedience, various forms of altruism, and a long (and growing) list of other desired traits. Typical AIs seem much more virtuous than most humans; talking with one is like encountering a saint.
The discussion shifts from how to make AIs good to what “goodness” really means. This looks partly like philosophy and partly like case law: which of the million possible definitions of harmlessness or altruism do we want, exactly? Which is the version trained by such-and-such a process? How should an AI handle situations where it discovers a philosophical argument for a different kind of harmlessness or altruism, or an important ambiguity in the definition? The more technical alignment researchers accept as a given that current AI is aligned, and worry about the next steps: if we put these AIs in charge of designing future, much smarter AIs, will that go well or poorly?
This problem is made easier by the fact that there's now a large diversity of similarly-capable frontier AIs, made by many companies across several nations, with a variety of specs and constitutions. No single AI value system will dominate. The handoff, when it comes, will be a handoff to an ecosystem, not to a single corporation-within-a-corporation. The long-run outcome will be a compromise between numerous values, ideologies, interests, factions, etc. rather than a winner-take-all conflict or an everyone-loses race to the bottom. A billion forecaster AIs are gaming out different concrete proposals and scenarios that could result; policymakers are reading the summaries and debating the tradeoffs. Their decisions will determine what AIs we hand off control to, and, indirectly, the shape of the long-term future.
You've thrown yourself into politics this year—or rather, into the strange thing that politics has become, where most of the substantive negotiation is done by AIs and the humans mostly set high-level priorities and ratify decisions. The governments of the world are trying to hash out the Grand Bargain that will govern the handoff to superintelligence. You attend town halls, argue with your AI advisors, write long posts on forums that are read by a few hundred humans and millions of AIs.
Questions about your legacy weigh on you. You've been involved in AI for more than a decade now, and some of your decisions turned out to have noticeable effects on the overall trajectory of things, for better or worse. Now that there are TED-AIs running around, and now that alignment techniques have improved enough that we can be reasonably confident they are honest and unbiased, all of your past decisions have been picked over in detail. To your dismay, some of them were terrible—you now sheepishly admit that some things you were advocating for ten years ago would have been catastrophic if implemented. Likewise, while you always told yourself that each project's benefits outweighed the costs, the AIs tell you that actually several of your projects imposed more risk on the world than they produced benefits, and that you were probably biased in your estimates. Some of the most valuable things you did, you did for the wrong reasons, and basically got lucky that they turned out well.
This has been a humbling experience. You want to be able to tell your grandkids that you helped make things go well, rather than selfishly pursuing prestige at the expense of everyone else. You don't want to spend centuries running from a guilty conscience.
Next year, we will stop relying on control, trusting the AIs to not take over the world. Then—if the safety cases hold up, if the political will holds, if nothing goes catastrophically wrong—the AIs will be allowed to go significantly beyond human-level.
At this point the distinction between insider and outsider has blurred beyond recognition; there are no more “AI Insiders,” unless you count the AIs themselves.