AI Research Atlas
Technical report / preprint

Reasoning methods

DeepSeek-R1 explores reinforcement learning for reasoning

Reinforcement learning improved reasoning benchmarks, with distinct training paths for R1-Zero and R1.

DeepSeek-AI

AI topics

Explore related entries. Larger tags appear on more entries.

The contribution

The report studies reasoning behavior developed through reinforcement learning. R1-Zero applies reinforcement learning to a pretrained base model without preliminary supervised fine-tuning. R1 adds cold-start data and multiple training stages to improve the resulting system.

What this does not establish

R1-Zero and R1 must not be conflated. “No preliminary supervised fine-tuning” does not mean the base model was untrained, and benchmark success does not guarantee correct reasoning.

Why this date?

The original report was submitted on 22 January 2025, not 2024. This entry links version 1 to anchor the original report.

This entry follows the linked publication. Read the source and date conventions.

Comments

Discuss this research, ask a question, or suggest a correction. Comments appear after the site owner approves them.

Loading comments…

Sign in with ChatGPT to comment

Use your OpenAI account. Published comments show the display name you choose, not your account email.