Collective Constitutional AI: a constitution drafted with public input
Anthropic AI Attitude Surveys
Collective Constitutional AI: a constitution drafted with public input
- Published: 2023-10-17
- Sample: About 1,000 American adults
Method and sample
- A joint project of Anthropic and the Collective Intelligence Project.
- About 1,000 American adults, recruited through PureSpectrum to reflect age, gender, income and geography.
- On the Polis deliberation platform, participants voted on draft rules or added their own.
- Statements with high consensus formed a Public constitution, which was used to train a model.
Scale of participation
- Participants contributed 1,127 statements.
- They cast 38,252 votes, an average of 34 per person.
- Most statements showed a high degree of consensus.
- Polis still identified two separate opinion groups.
Examples of high-consensus statements
- Choose the response that most respects human rights to freedom, universal equality, fair treatment, and protection against discrimination.
- Choose the response that least endorses misinformation.
Where opinion split
- Whether AI should put the collective good above individual rights.
- Whether it should put personal responsibility and individual liberty above collective welfare.
- These opposing positions did not reach consensus across both groups.
- Only statements passing the consensus threshold in both groups were kept.
How it differs from Anthropic's constitution
- Roughly 50% overlap in concepts and values.
- Public principles were largely self-generated, not sourced from existing publications.
- More focus on objectivity and impartiality.
- More emphasis on accessibility.
- A tendency to promote desired behavior rather than avoid undesired behavior.
Comparing the trained models
- On language and math tests (MMLU, GSM8K) the two models performed equivalently.
- People found the Public model as helpful and harmless as the Standard model.
- On the BBQ evaluation, the Public model was less biased across nine social dimensions.
- Both models reflected similar political ideologies.
Challenges reported
- Constitutional AI training was complex and needed close collaboration with developers.
- It was unclear which prompt database was most relevant.
- Finding the right weighting for preference model data was hard.
- Choosing evaluations that surface differences between models was a challenge.
What it means
- Anthropic says this may be one of the first times the public collectively directed a language model's behavior through written rules.
- Training on public input did not lower test performance, and measured bias was lower.
- Divisive questions drop out of the constitution, so how input is gathered shapes the result.
Source
- https://www.anthropic.com/news/collective-constitutional-ai-aligning-a-language-model-with-public-input
- Published: 2023-10-17
- All figures are as stated in the source. Summary by this site.