AI Research Atlas
Technical report

Efficient models

Mixtral activates a sparse set of experts

Sparse routing selected two of eight feedforward experts per token at each layer.

Albert Q. Jiang and colleagues

AI topics

Explore related entries. Larger tags appear on more entries.

The contribution

Mixtral uses a sparse mixture-of-experts architecture: a routing mechanism chooses which expert networks process each token. This separates total model capacity from the amount of expert computation used for an individual token.

What this does not establish

Only part of the model is active per token, but storing the model still requires its full weights. Sparse activation is not the same as a small total model.

Why this date?

The paper was submitted on 8 January 2024. The model was released in December 2023, so both years can be correct when clearly labeled.

This entry follows the linked publication. Read the source and date conventions.

Comments

Discuss this research, ask a question, or suggest a correction. Comments appear after the site owner approves them.

Loading comments…

Sign in with ChatGPT to comment

Use your OpenAI account. Published comments show the display name you choose, not your account email.