Skip to content

A hard question

Could alignment and better oversight solve this?

Alignment research could reduce risk. The disputed step is relying on it to justify systems whose failures could exceed our ability to recover.

2 min readArgument & objectionUpdated 23 September 2026

The strongest objection

AI systems can be trained, evaluated, monitored, and redesigned. Better interpretability, robust control methods, and carefully structured deployment could make more capable systems safer. Greater capability might even help identify weaknesses and improve oversight. It would be unjustified to assume that every possible technical solution must fail.

What the prevention argument asks

Success on observed tests does not automatically establish reliable behavior in unfamiliar situations. The assurance problem becomes harder when a system can understand its evaluation, plan over longer periods, or exploit dependencies outside the test environment. The relevant standard must concern the actual deployed system, including its tools, permissions, and interactions with other systems.

Using another AI as an evaluator may help, but it also creates a question about the evaluator's reliability and independence. Adding more oversight components does not establish that their shared blind spots have been removed.

There is also a political question. A system that faithfully follows one owner's instructions can still reduce everyone else's power. Technical obedience and a legitimate distribution of authority are different requirements.

The hinge of disagreement

How much assurance is achievable before deployment, and how severe are the consequences if that assurance is wrong? W360's position is that potentially irreversible loss of control requires a much stronger justification than improved benchmark behavior. A demonstrated method that survives independent adversarial testing at the relevant capabilities would materially strengthen the case for revising restrictions.

Keep following the question

Read next.