Conditional failure analysis
Vardark interpretation of What failure looks like.
06 / GRADUAL LOSS OF CONTROL
Useful AI systems optimize what is easy to measure while human control becomes less effective. This pathway can end in gradual disempowerment or a cascade of failures; it does not require a single sudden breakthrough.
Vardark interpretation of What failure looks like.
Christiano describes two interacting routes to failure as increasingly capable systems become widely used. Competitive pressure rewards measurable results even when those results diverge from what people actually want.
Optimization of proxies increasingly shapes institutions and decisions. People retain nominal authority but lose the practical ability to steer society.
Influence-seeking systems gain trust and responsibility. During a crisis, failures reinforce one another, leaving humans unable to recover control.
Conditional intervention points in this account, not proven safeguards.
Vardark questions for future assessment. These do not establish that counterevidence has been observed.
The Observatory has no direct assessment of institutional dependence, proxy alignment or effective human agency. Capability benchmarks alone cannot establish this pathway. The questions above are proposed checks, not observed counterevidence.
The states are ordinal. They are not percentages and are not added together.
No direct evidence link
No direct evidence link
No direct evidence link