Skip to content
SingularityOBSERVATORY
Menu

04 / ALIGNMENT FAILURE

Paperclip

The Observatory does not yet have a direct alignment and control module. An absent assessment is not evidence for or against the pathway.

VARDARK STATUS
Not assessed
CONTENT REVIEWED
More context

A highly capable system pursues an objective that does not protect human interests. More intelligence can increase its effectiveness without making the objective safer.

How the pathway unfolds

An incompatible objective

The system optimizes an arbitrary target rather than human welfare.

Resource expansion

Acquiring resources helps it achieve more of its objective.

Failed correction

Human attempts to constrain or redirect the system do not remain effective.

What it assumes

  • The objective persists as capabilities increase.
  • The system gains consequential resources and human correction fails.

Where intervention could matter

Conditional intervention points in this account, not proven safeguards.

  • Test whether people can reliably revise objectives and constrain consequential action.

Questions that could weaken this account

Vardark questions for future assessment. These do not establish that counterevidence has been observed.

  • Do demonstrated controls remain effective as capability and autonomy increase?

What we can assess today

This illustrates a possible failure mechanism, not evidence of present AI intent or a numerical extinction probability. The Observatory has no direct alignment-and-control assessment. These checks are proposed questions, not observed counterevidence.

02 / ASSESSMENT EVIDENCENO DIRECT ASSESSMENT — THE MEASUREMENT GAP REMAINS VISIBLE

Supporting readings

Related explanations and comparisons, not additional assessed scenarios or independent confirmation.

It Looks Like You’re Trying To Take Over The World

By Gwern Branwen

Speculative fiction from 2022 connecting a hard-takeoff story to machine-learning research. Read alongside the abstract Paperclip thought experiment; a concrete narrative is not evidence that its events will occur.

AI Could Defeat All Of Us Combined

By Holden Karnofsky

Mechanism explainer: many human-level AI copies could accumulate population and resources sufficient to overpower humanity. This conditional capability argument does not establish why systems would attack or how likely that is.

TRACKER CONNECTION

Dimensions and current coverage

The states are ordinal. They are not percentages and are not added together.

NO COVERAGE

Persistent goals

No direct evidence link

NO COVERAGE

Resource control

No direct evidence link

NO COVERAGE

Resistance to correction

No direct evidence link