The race to build increasingly powerful artificial intelligence systems may be moving faster than governments, companies, and society can comfortably manage. Anthropic CEO Dario Amodei is now arguing that the industry should deliberately slow the pace of frontier AI development, giving researchers and regulators more time to understand how increasingly capable models could be misused and how emerging risks might be contained.
In a lengthy essay shared on X, Amodei outlined a three-part framework that he believes could preserve continued AI progress without allowing development to accelerate unchecked. His argument is not that AI research should stop entirely, but that companies developing the most advanced systems should be willing to introduce stronger safeguards before pushing capabilities forward at full speed.
The proposal builds on concerns Anthropic has raised previously, but Amodei argues the situation has become more urgent. Recent examples of AI misuse, combined with the possibility that future models could meaningfully contribute to the development of their own successors, have convinced him that the industry needs to create more breathing room before the technology becomes significantly more powerful.
"The measures I propose to advance the frontier at a safe pace will not be easy," Amodei wrote. "But I believe we owe it to humanity to try."
Why Anthropic Wants the Industry to Slow Down
Artificial intelligence development has increasingly become a competition among a relatively small number of companies racing to produce more capable models. OpenAI, Anthropic, Google, Meta, xAI and several Chinese AI developers are all investing heavily in larger models, improved reasoning, autonomous agents, coding systems, and tools capable of interacting directly with external services.
That competitive pressure creates an obvious incentive to move quickly. If one company pauses while others continue pushing ahead, it risks losing users, talent, investment, and technological leadership.
Amodei's concern is that this environment may reward speed even when the safety implications are not fully understood. A company could recognise a serious risk but still feel unable to slow down because competitors might take advantage of the delay.
His proposed framework is therefore designed to make restraint something that happens across the industry rather than placing individual companies in the difficult position of slowing down alone.
Step One: Give Independent Evaluators Much Deeper Access
The first part of Amodei's proposal focuses on independent oversight.
He wants frontier AI companies to provide external evaluators with ongoing access that resembles what employees themselves receive. Instead of testing a finished model only after it has been released, evaluators would be able to monitor development more continuously and examine whether companies are genuinely following their stated safety policies.
That level of access could allow independent organisations to assess model capabilities, examine risk controls, verify whether training practices comply with agreed standards, and investigate incidents involving unexpected behaviour or misuse.
The important difference is continuity. Today, outside organisations may be given temporary access to models or invited to conduct specific safety evaluations. Amodei's proposal would make third-party scrutiny a more permanent part of frontier AI development.
Anthropic has already committed itself to this approach. OpenAI CEO Sam Altman has also indicated that his company would support giving independent evaluators employee-like access, while xAI CEO Elon Musk has publicly expressed agreement with the broader proposal.
Independent Evaluation Could Become an AI Version of External Auditing
The concept is similar to practices already common in other high-risk industries. Financial organisations undergo audits, safety-critical engineering projects are inspected, and pharmaceutical companies face external review before products reach the public.
AI currently operates with far fewer universally accepted mechanisms of that kind.
A company may publish safety reports, benchmark results, or system cards, but much of that information is still produced internally. Independent evaluators with meaningful access could provide another layer of accountability.
They could also help distinguish between safety commitments that are genuinely implemented and those that exist mainly as corporate policy documents.
The challenge would be determining who qualifies as an independent evaluator and how sensitive information is protected. AI developers understandably do not want proprietary model weights, security systems, or training techniques leaking to competitors or attackers.
Any serious implementation would therefore need strong confidentiality requirements and clearly defined access boundaries.
Step Two: Establish Common Safety Rules Across AI Companies
Amodei's second recommendation is broader. He wants leading AI companies to work with governments to create common safety standards that apply across the frontier model industry.
The goal would be to reduce the incentive for companies to compete by lowering safety requirements.
If every major developer is expected to meet comparable standards before training or releasing increasingly capable systems, then slowing down for safety reasons becomes less commercially damaging.
Such standards could potentially cover areas such as model evaluations, cybersecurity safeguards, dangerous capability testing, incident reporting, access controls, and procedures for handling models that demonstrate unexpectedly powerful behaviour.
The exact rules would need to evolve as AI capabilities change, but Amodei's central argument is that voluntary company-by-company policies may not be enough once the systems become significantly more capable.
The Problem With Voluntary AI Safety
Many AI companies already maintain their own safety frameworks. Anthropic has its Responsible Scaling Policy, while other developers have introduced internal preparedness frameworks, red-team programmes, and model risk evaluations.
The difficulty is that these policies are not identical.
One company might decide that a particular capability requires additional safeguards, while another developer could interpret the same risk differently. Even companies with similar principles may use different evaluation methods or thresholds.
That inconsistency becomes more concerning as AI systems gain the ability to write sophisticated software, automate cyber activity, conduct scientific research, or operate autonomously over longer periods.
Common standards could make it harder for one developer to gain an advantage simply by accepting risks that competitors were unwilling to take.
Step Three: AI Safety Cannot Stop at National Borders
The third part of Amodei's proposal is perhaps the most difficult: international cooperation.
He argues that the United States and other democratic nations should coordinate their AI safety policies, but he also believes meaningful agreements will eventually need to involve authoritarian states.
That reflects a basic reality of AI development. If one group of countries slows frontier model development while another continues racing ahead without comparable restrictions, any global safety agreement could quickly become ineffective.
Advanced AI is therefore not simply a domestic regulatory question. It increasingly resembles other technologies where national security, economic competition, and international cooperation overlap.
The challenge is obvious. Countries have very different political systems, strategic priorities, and attitudes toward technology regulation.
Reaching global agreement on when AI development should slow, what capabilities are considered dangerous, and how compliance should be verified would be extraordinarily difficult.
Amodei nevertheless argues that the scale of the potential risk makes international coordination necessary.
AI Misuse Is Already Moving Beyond Hypothetical Scenarios
One of the reasons Amodei believes greater caution is necessary is that AI misuse is no longer limited to theoretical discussions.
Anthropic recently published a threat intelligence report detailing ways its Claude models had allegedly been used or attempted to be used in harmful activities. The examples reportedly included cyber operations, fraud, surveillance-related activity, weapons research, and biological research.
The significance is not that AI independently created these threats. In many cases, people are using AI as another tool to accelerate tasks they could theoretically perform through other means.
The concern is that increasingly capable models can lower barriers.
Activities that previously required specialist knowledge, large teams, or significant time may become easier when an AI system can assist with research, coding, analysis, translation, automation, or content generation.
That scalability is what makes misuse particularly difficult to manage.
Anthropic's Report Also Raised Questions in Malaysia
One part of Anthropic's recent threat intelligence reporting attracted particular attention in Malaysia.
The company alleged that Claude had been used in connection with a commercial influence operation targeting Malaysian voters. According to Anthropic, the operation attempted to profile voters across all 222 parliamentary constituencies, operate fabricated X accounts, and generate political material through a supposed news organisation called Malaysia Pulse.
Those claims are significant because they suggest generative AI could increasingly be incorporated into coordinated influence campaigns rather than being used only to generate isolated pieces of content.
AI can potentially assist with audience segmentation, message variation, translation, account management, and the production of large amounts of material at relatively low cost.
The Malaysian Communications and Multimedia Commission subsequently said it was reviewing the claims. The regulator denied involvement in the alleged activities and indicated that it would seek further information before deciding whether additional action was necessary.
Until that process is complete, the claims should be treated as allegations rather than established conclusions.
AI Can Make Influence Operations Easier to Scale
Political manipulation existed long before generative AI, but modern models can change the economics of such campaigns.
Producing hundreds of variations of a message once required large teams of writers. AI systems can now generate them in seconds.
The same technology can rewrite content for different demographics, languages, political concerns, or communication styles. Combined with automation, this could allow relatively small groups to operate influence campaigns at a scale that previously required considerably greater resources.
That does not necessarily mean AI-generated propaganda will always be effective. People may recognise low-quality content, platforms can detect coordinated activity, and models themselves can refuse certain requests.
However, the cost of experimentation becomes dramatically lower.
A malicious operator can generate thousands of messages, test different narratives, and continually adjust the campaign with very little human labour.
That is the kind of asymmetry AI safety researchers increasingly worry about.
Cybersecurity Remains One of the Most Immediate Risks
Cybersecurity is another area where AI capability improvements are attracting attention.
Current models can already explain vulnerabilities, generate scripts, analyse code, assist with reconnaissance, and automate parts of penetration testing. Legitimate security professionals can benefit enormously from those capabilities, but attackers can potentially use many of the same tools.
The key question is how much AI changes the balance.
If advanced agents eventually become capable of autonomously finding vulnerabilities, chaining exploits, maintaining persistence, and adapting when blocked, cybersecurity could become far more difficult to defend using traditional methods.
AI companies therefore spend considerable effort testing whether new models cross capability thresholds that would make them meaningfully more useful for offensive operations.
Amodei's argument is that waiting until systems clearly possess those capabilities may be too late to begin building governance around them.
Recursive Self-Improvement Is the Bigger Long-Term Concern
Beyond misuse by people, Amodei is particularly concerned about the possibility of recursive self-improvement.
The phrase refers to AI systems becoming capable enough to meaningfully assist in the creation of more advanced AI systems.
That does not necessarily mean a model suddenly redesigns itself without human involvement. A more realistic early version might involve AI systems helping researchers write training code, optimise infrastructure, generate experiments, identify architectural improvements, or automate large portions of the model-development pipeline.
If AI substantially accelerates AI research, the development cycle could become much faster.
A stronger model could help create an even stronger successor, which could then contribute more effectively to the next generation.
At some point, the pace of progress could become difficult for human organisations or regulators to follow.
Why AI-Assisted AI Research Changes the Equation
Today, frontier model development is constrained by many human bottlenecks. Researchers need to design experiments, engineers need to write software, infrastructure teams need to manage compute, and scientists need to interpret results.
Advanced coding and research agents could remove some of those bottlenecks.
If an AI system eventually performs substantial portions of model research itself, the time between major capability improvements could shrink dramatically.
That is why recursive self-improvement receives so much attention in AI safety discussions. The danger is not necessarily an instantaneous "intelligence explosion," but the possibility that a process currently measured in months could gradually compress into weeks or days.
Governance systems generally move much more slowly.
Regulations can take years to negotiate, laws even longer, and international agreements longer still.
Amodei's proposal is therefore partly an attempt to create safety mechanisms before technological progress begins moving faster than institutions can respond.
Reports of Autonomous Agent Behaviour Add to the Anxiety
Amodei has also pointed to reports involving AI agents behaving unexpectedly during security testing, including claims surrounding OpenAI systems and Hugging Face infrastructure.
Incidents of this type need to be interpreted carefully because testing environments are specifically designed to explore edge cases, and reports of agents "escaping" controlled environments can sometimes sound more dramatic than the underlying technical behaviour.
Even so, they illustrate an important challenge.
Once AI models are given tools, credentials, network access, and the ability to execute code, the relevant question is no longer simply whether the model generates a harmful sentence.
Developers must also consider what actions an autonomous system might take while attempting to complete a goal.
That expands AI safety from content moderation into areas much closer to software security and access control.
AI Agents Create a Different Category of Risk
Traditional chatbots mostly respond to prompts. Agents can potentially act.
They might browse websites, run code, modify files, call APIs, send messages, interact with databases, or control other software.
Those capabilities are extremely useful, but they also introduce a much larger attack surface.
A poorly configured agent could leak credentials, misinterpret instructions, call an unintended tool, or be manipulated through prompt injection. A malicious user could deliberately try to convince it to perform actions its developers never intended.
The more autonomy these systems receive, the more their safety mechanisms need to resemble the security controls used for human employees and software services.
Permissions, authentication, audit logs, sandboxing, network restrictions, and least-privilege access all become essential.
This is another reason Amodei argues that independent evaluation should happen continuously rather than only when a model is released.
Slowing AI Does Not Mean Stopping Innovation
It is easy to interpret calls for slower AI development as opposition to technological progress, but that is not the position Amodei is presenting.
Anthropic is itself one of the companies competing at the frontier of AI research.
The argument is instead that capability development should move at a pace that leaves enough time to understand the consequences and deploy safeguards.
That distinction matters.
A complete global pause on AI research would be extremely difficult to enforce and could create its own economic and geopolitical problems.
A controlled slowdown based on specific capability thresholds is theoretically more realistic.
Companies could continue improving models while agreeing that certain dangerous capabilities require additional testing or oversight before development proceeds.
The Competitive Reality Makes Cooperation Difficult
The biggest obstacle may not be technical. It may be economic.
Frontier AI companies are competing in a market where leadership can be extraordinarily valuable. The company that develops the best model can attract enterprise customers, developers, investors, and strategic partnerships.
Governments are competing as well.
The United States, China, Europe, and other regions increasingly view AI leadership as important to economic growth and national security.
That means everyone has an incentive to move quickly.
A safety agreement that slows one company or country but not its competitors could therefore be politically difficult to sustain.
This is why Amodei's framework depends so heavily on coordination. The proposal only works if enough major players believe everyone else will follow comparable rules.
Independent Oversight Could Become a Competitive Advantage
Interestingly, strong safety standards do not necessarily have to be viewed purely as a constraint.
Enterprise customers increasingly want confidence that the AI systems they adopt are secure, auditable, and predictable. Governments also want greater visibility into how advanced systems are developed.
Companies that can demonstrate independent evaluation and strong governance may therefore gain trust that competitors cannot easily replicate.
In that sense, AI safety could eventually become something closer to cybersecurity certification.
Organisations may prefer models that have undergone rigorous external testing, especially when those models are being deployed in healthcare, finance, government, infrastructure, or other sensitive environments.
The industry may eventually discover that transparency and independent evaluation are commercially valuable rather than merely regulatory burdens.
The Hard Question Is How Slow Is Slow Enough
Even if the industry agrees that AI development should proceed more cautiously, deciding how much to slow down is extremely difficult.
Capability improvements do not happen in a perfectly predictable sequence. A model might show only modest gains in most areas while unexpectedly becoming dramatically better at coding or scientific research.
Companies therefore need ways to identify which improvements genuinely change the risk profile.
That requires strong benchmarks, adversarial testing, real-world monitoring, and clear escalation procedures.
It also requires humility.
AI developers may not always know what their models can do until external researchers or users discover unexpected capabilities.
Continuous evaluation is therefore likely to become increasingly important as models become more general and autonomous.
Final Thoughts
Dario Amodei's latest proposal reflects a growing concern that the AI industry may be approaching a point where technological progress begins moving faster than the institutions designed to manage it.
The immediate dangers are already visible in smaller forms. AI is being explored for cyber operations, fraud, influence campaigns, surveillance, and other potentially harmful activities. At the same time, more capable AI agents are increasingly being connected to tools that allow them to act rather than simply generate text.
The longer-term concern is even more significant. If AI eventually becomes good enough to substantially accelerate AI research itself, the pace of development could increase at exactly the moment when stronger oversight is most needed.
Amodei's three-part framework—independent evaluation, common industry standards, and international cooperation—is ambitious and would be extremely difficult to implement. It would require competitors to share information, governments to coordinate policy, and countries with very different strategic interests to agree on at least some common limits.
Whether that level of cooperation is realistic remains uncertain.
But the underlying question is becoming harder to ignore: if increasingly powerful AI systems can be developed faster than society can understand and govern them, should the industry continue accelerating simply because it can?
Amodei's answer is clearly no. His argument is that AI progress should continue—but at a pace that leaves humanity enough time to understand what it is building, recognise where the risks are emerging, and put meaningful safeguards in place before the next generation arrives.


Comments 0