Aligning AI to values

What models should be trained toward, who supplies that target, and how values are elicited from real people.

Curated by
TBD
Relevant for
10 grid cellsAgent-agent vows and 9 others
Write to
values@agi-institutions.orgAn operator routes it to the right experts.

Selected papers

3

Must-reads for understanding the field and why it matters for AGI institutions.

  1. Argues that aligning powerful AI requires institutions able to represent and revise thick, context-sensitive human values, not only individual preference aggregation.

  2. Shows why preference satisfaction is too thin a target for alignment and develops a broader account of what AI systems should respond to.

  3. Elicits the considerations behind people's choices rather than ratings, and reconciles them across a population into a moral graph that can serve as an alignment target.

Work in the field

13

Recent work in the field worth knowing.

Foundations

3

Older work that gives background on the field.

This list isn't exhaustive. About the lists.