AI Research Atlas
Preprint

Alignment

Principles guide AI feedback

Written principles guided self-critique and AI-generated feedback during training.

Yuntao Bai and colleagues

AI topics

Explore related entries. Larger tags appear on more entries.

The contribution

Constitutional AI used a set of principles to guide revisions of model responses and to help generate preference feedback. The paper explored reducing dependence on human labels for judgments of harmfulness while improving the behavior of an assistant.

What this does not establish

Human choices remain in the principles and other parts of the training process. The method does not establish that the model is harmless in every situation.

Why this date?

The first arXiv submission was 15 December 2022.

This entry follows the linked publication. Read the source and date conventions.

Comments

Discuss this research, ask a question, or suggest a correction. Comments appear after the site owner approves them.

Loading comments…

Sign in with ChatGPT to comment

Use your OpenAI account. Published comments show the display name you choose, not your account email.