Thought experiment about objectives and control
Vardark interpretation of Ethical Issues in Advanced Artificial Intelligence.
04 / ALIGNMENT FAILURE
The Observatory does not yet have a direct alignment and control module. An absent assessment is not evidence for or against the pathway.
Vardark interpretation of Ethical Issues in Advanced Artificial Intelligence.
A highly capable system pursues an objective that does not protect human interests. More intelligence can increase its effectiveness without making the objective safer.
The system optimizes an arbitrary target rather than human welfare.
Acquiring resources helps it achieve more of its objective.
Human attempts to constrain or redirect the system do not remain effective.
Conditional intervention points in this account, not proven safeguards.
Vardark questions for future assessment. These do not establish that counterevidence has been observed.
This illustrates a possible failure mechanism, not evidence of present AI intent or a numerical extinction probability. The Observatory has no direct alignment-and-control assessment. These checks are proposed questions, not observed counterevidence.
Related explanations and comparisons, not additional assessed scenarios or independent confirmation.
By Gwern Branwen
Speculative fiction from 2022 connecting a hard-takeoff story to machine-learning research. Read alongside the abstract Paperclip thought experiment; a concrete narrative is not evidence that its events will occur.
By Holden Karnofsky
Mechanism explainer: many human-level AI copies could accumulate population and resources sufficient to overpower humanity. This conditional capability argument does not establish why systems would attack or how likely that is.
The states are ordinal. They are not percentages and are not added together.
No direct evidence link
No direct evidence link
No direct evidence link