What’s the big risk here? 

Years ago, many hypothesised that governments would reign supreme in producing technologies symbolising power and authority. Such technology grew to be feared for their nigh-catastrophic potential to be destructive. Today, we have such technology being born and bred on servers, having been developed in office towers. 

Artificial intelligence is growing to be more than a source of entertainment and convenience, it is able to instruct one to build bombs and phish. For this reason, AI companies like Anthropic anticipated the need for safety precautions and created Responsible Scaling Policies or RSPs, to ensure they take the very necessary first-steps to curb harmful risks. One reason RSPs were readily produced as standards was due to increasing literature on CBRN-led risks, which could be exponentially dangerous if left unscrutinised. As per METR 2023, the creation of RSPs are an extremely necessary one, given “biologists have argued that large language models (LLMs) might remove some barriers encountered by historical biological weapons efforts… Today, someone with enough biology expertise could likely build a dangerous bioweapon and potentially cause a pandemic.”

Why are RSPs’ safety levels important?

Responsible Safety Policies, refer to the policy framework that Anthropic can use in order to scale their AI while aligning with ethical safety standards. RSPs are a wavemaker in terms of self-regulation on the part of AI giants, with many following suit. In 2023, this public commitment was first-of-its-kind and included 3 main goals. Anthropic aimed to ensure its risk governance was proportional, iterative and exportable. Proportional to ensure they can balance innovation with safety, iterative to ensure experimentation was realistic in light of frontier model’s rapid and unyielding progress, and exportable to provide a proof of concept that could develop into a readily adopted and respectable industry standard for self-governance. One could say that since 2023, they have succeeded in most aims. 

To further ensure these principles are upheld, RSP documents encased several sections on its AI Safety Level Standards or ASL which refer to levels of protections accorded based on the:

  1. Identified capability threshold
  2. Required safeguards for said identified threshold
  3. Any changes to identified category post-capability assessment

Here is a brief summary of the ASLs and the capability classifications mapped to each, as per Anthropic, 2023: 

  1. ASL-1 refers to systems which pose no meaningful catastrophic risk, for example a 2018 LLM or an AI system that only plays chess.
  2. ASL-2 refers to systems that show early signs of dangerous capabilities – for example ability to give instructions on how to build bioweapons – but where the information is not yet useful due to insufficient reliability or not providing information that e.g. a search engine couldn’t. Current LLMs, including Claude, appear to be ASL-2.
  3. (Where we are now) ASL-3 refers to systems that substantially increase the risk of catastrophic misuse compared to non-AI baselines (e.g. search engines or textbooks) OR that show low-level autonomous capabilities.
  4. ASL-4 and higher (ASL-5+) is not yet defined as it is too far from present systems, but will likely involve qualitative escalations in catastrophic misuse potential and autonomy.

RSPs clearly dictate the safeguard assessments in place for each ASL and the decision-making framework behind giving the go-ahead to deploy a model. Ultimately, the model must be distance in ability from existing thresholds, or have surpassed them but under the stipulated ASL-2/3 safeguards with routine assessments. Results of core system tests are stated in System Cards such as this one

ASLs also comprise of 2 key components, one being the deployment standard and the other being the security standard. The former consists of acceptable use model cards, reporting channels and automated detection of harmful requests. The latter sets the precedent for security system designs meant to thwart malicious actors hoping to manipulate the models for their own purposes. Accessing model weights in particular, is an outcome that takes priority when devising such measures. If one circumvents security measures, accessing model weights are what allow malicious actors to utilise a model without its deployment protections; thus making the ASL safeguards benign to intruders. 

In line with Anthropic’s 3rd goal, OpenAI has also made their own version of ASLs within their policy framework  named the Preparedness Framework. While OpenAI’s framework aims to engage in preparation for similar risk categories, Anthropic looks at RSPs on a more holistic level, with a focus on CBRN risks, autonomous R&D capabilities and cyber operations. 

In April 2025, OpenAI began to look deeply into “research categories”, which are risk categories that pose serious harm but currently lack mature threat models. Perhaps this category will be the subject of speculation and research for the next few sets of AI safety standards, especially in terms of issues such as model sandbagging (models intentionally underperforming), models engaging in autonomous replication and adaptation (goal-drifting). 

Being able to deal with such threats on the horizon would mean the need for ever-evolving standards such as Anthropic’s RSPs, and perhaps its widespread adoption of some form in companies elsewhere producing and tuning AI systems. 

What is Claude Opus 4? Why is it in the news? 

Claude Opus 4 is Anthropic’s most advanced large language model to date, designed to handle prompts that require complex reasoning and problem solving. Some significant updates that it has in contrast to its predecessor are enhanced memory from previous prompts and higher stamina for long duration prompts. Probably the biggest update that it has is its coding performance: the model got a 72.5% on the SWE-bench, which is the benchmark test for engineering tasks. That is a significant jump from OpenAI’s GPT-4.1, which only achieved 54.6%.

In short, Claude just got a huge upgrade and is now one of the most powerful models on the market. However, as AI systems like Claude become more powerful, they become more prone to unintended behaviors. That’s where the ASLs come in. Claude Opus 4 was in fact so significantly able, that the 3rd tier of ASLs were activated, a standard that was previously reserved for being adequate to handle CBRN risks and some level of autonomous AI research activities.

What is the significance of activating Level 3? 

You can reasonably assume that a prompt triggering Safety Level 3 protection would mean that it poses a great amount of risk in terms of ensuring its deployment, or in terms of ensuring its security in light of model capabilities.

In this case, the model is within ASL-3 threshold and can in fact, answer questions relating to building bio-weapons or bombs. Since AI could potentially help a user create CBRN weapons and mass destruction, Level 3 aims to mitigate this on both the deployment and security fronts. However, in order to limit access to this type of sensitive data, it would require rigorous capability assessments and routine monitoring to ensure appropriate deployment is done for models capable of providing such information. Security measures would have to be adequate to ensure no actors can access the model to get such information via rote prompting or through cyber attacks.

Anthropic has ensured that such threats are covered in the security practices, such that the model cannot respond to such queries and there are “more than a 100 different security controls” that involve preventive measures and sharp, unrelenting monitoring controls to keep non-state actors from accessing model weights.

The emphasis is on the term “non-state” as there is little one could do to keep the likes of state actors from accessing such information. This could prove to be a grave issue moving forward as models keep progressing and the discussion of AI has a political edge to it with OpenAI’s discussion of “democratic AI” and frontier AI companies’ foray into political theatre.

However, perhaps the activation of ASL-3 is not all that bad. If (and when) triggered, Level 3 could also provide relevant data to Anthropic regarding what they might have to re-train or update in order to best interpret a prompt. This could be a way the private company experiments with its future safety levels.  Deploying a model with hesitancy but then testing the appropriateness of this decision would be one way to determine if Level 4, 5 and perhaps 6 are necessary to properly ideate, conceptualise, and roll out. Many researchers around the globe await with a bated breath to see the results of the deployment.

Should we continue deploying models at this frequency? What checks and balances do we need beyond RSPs, if any at all?

The activation of an additional safety level is not a crisis, but rather an announcement that an additional layer of security is needed. It shows that Anthropic is taking time to address risks that many other companies are ignoring, and that a conscious effort is being made to minimise the risks of their progress. Rather than focusing only on profit, Anthropic is choosing to remain transparent and maintain human-in-the-loop practices, with their governance documentation consistently growing.

This development signals yet again just how fast AI is growing. Artificial Intelligence is no longer just a tool to get the dirty work done—it's starting to become a full-fledged partner that can reason with potentially minimum  ethical repercussions.

However, relying on safety levels alone is not a long-term solution. As we continue to use AI in law, education, and healthcare, the risks will only continue to grow. Just adding extra safety levels is not enough to ensure accountability without some oversight. At the end of the day, or at least for the near future, humans will always be needed to ensure Claude and the models that come after are regulated sufficiently.

There is much more that Anthropic and other companies could do. Perhaps instead of general categories within ASL-3 or research categories that include potential risk outcomes like self-adaptation, companies should specify different categories for a more granular set of risk types at a given safety level. This could mean elaboration on types of misuse, a range of self-adaptation scenarios, sectoral and systemic risks  (e.g. financial market manipulation, access to public records). 

More importantly, IAPS’s Bill Anderson-Samways suggests that Anthropic and other companies ought to outrightly declare when they will alert government authorities of identified risks as part of their accountability measures. The author rightly mentioned that at present, RSP documentation lacks any mention of mandating communication with governments outside of a narrow case in catastrophic risk management that mandates communication during Anthropic’s response to a bad actor scaling extremely quickly. The author suggests Anthropic ought to make such a commitment in line with the assigning of a capability threshold, like ASL-3 or 4. Governments should also be allowed to audit and aid capability and security safeguards-setting and other assessments, especially in the case of suspected ASL-3 and above. It is loud and clear– self-regulation cannot be all that stands in the way of normalcy and imminent disasters, with governments playing catch-up. 

All in all, Anthropic’s launch of ASL-3 is a step in the right direction. However, human intervention is still needed. After all, the true test of AI isn’t just how well it performs— it’s how well it behaves when no one is watching.