AI Research Atlas
Preprint / conference paper

Alignment

Human feedback shapes instruction following

Supervised examples and human preference feedback improved how language models followed instructions.

Long Ouyang and colleagues

AI topics

Explore related entries. Larger tags appear on more entries.

The contribution

InstructGPT combined supervised fine-tuning with a reward model trained from human rankings and reinforcement learning. Evaluators preferred outputs from the resulting models over those from the larger baseline in the studied comparisons.

What this does not establish

Preference improvements do not eliminate hallucinations, bias, or unsafe behavior. Results depend in part on the instructions and people providing feedback.

Why this date?

The first arXiv submission was 4 March 2022.

This entry follows the linked publication. Read the source and date conventions.

Comments

Discuss this research, ask a question, or suggest a correction. Comments appear after the site owner approves them.

Loading comments…

Sign in with ChatGPT to comment

Use your OpenAI account. Published comments show the display name you choose, not your account email.