Collective Constitutional AI: a constitution drafted with public input

Anthropic AI Attitude Surveys

Collective Constitutional AI: a constitution drafted with public input

  • Published: 2023-10-17
  • Sample: About 1,000 American adults

Method and sample

  • A joint project of Anthropic and the Collective Intelligence Project.
  • About 1,000 American adults, recruited through PureSpectrum to reflect age, gender, income and geography.
  • On the Polis deliberation platform, participants voted on draft rules or added their own.
  • Statements with high consensus formed a Public constitution, which was used to train a model.

Scale of participation

  • Participants contributed 1,127 statements.
  • They cast 38,252 votes, an average of 34 per person.
  • Most statements showed a high degree of consensus.
  • Polis still identified two separate opinion groups.

Examples of high-consensus statements

  • Choose the response that most respects human rights to freedom, universal equality, fair treatment, and protection against discrimination.
  • Choose the response that least endorses misinformation.

Where opinion split

  • Whether AI should put the collective good above individual rights.
  • Whether it should put personal responsibility and individual liberty above collective welfare.
  • These opposing positions did not reach consensus across both groups.
  • Only statements passing the consensus threshold in both groups were kept.

How it differs from Anthropic's constitution

  • Roughly 50% overlap in concepts and values.
  • Public principles were largely self-generated, not sourced from existing publications.
  • More focus on objectivity and impartiality.
  • More emphasis on accessibility.
  • A tendency to promote desired behavior rather than avoid undesired behavior.

Comparing the trained models

  • On language and math tests (MMLU, GSM8K) the two models performed equivalently.
  • People found the Public model as helpful and harmless as the Standard model.
  • On the BBQ evaluation, the Public model was less biased across nine social dimensions.
  • Both models reflected similar political ideologies.

Challenges reported

  • Constitutional AI training was complex and needed close collaboration with developers.
  • It was unclear which prompt database was most relevant.
  • Finding the right weighting for preference model data was hard.
  • Choosing evaluations that surface differences between models was a challenge.

What it means

  • Anthropic says this may be one of the first times the public collectively directed a language model's behavior through written rules.
  • Training on public input did not lower test performance, and measured bias was lower.
  • Divisive questions drop out of the constitution, so how input is gathered shapes the result.

Source

Back to list