AI Research Atlas
Technical report

Multimodal models

Gemini 1.5 studies long multimodal context

Long-context experiments included text, video, and audio within a multimodal model family.

Gemini Team

AI topics

Explore related entries. Larger tags appear on more entries.

The contribution

The report evaluates models designed to process very long inputs and retrieve or reason over information within them. Its experiments explore how extended context can support work with long documents and other media.

What this does not establish

Context capacity and retrieval results do not guarantee perfect comprehension of every long input. Claims remain tied to the report and its evaluation conditions.

Why this date?

The first arXiv submission was 8 March 2024. The record also contains later revisions.

This entry follows the linked publication. Read the source and date conventions.

Comments

Discuss this research, ask a question, or suggest a correction. Comments appear after the site owner approves them.

Loading comments…

Sign in with ChatGPT to comment

Use your OpenAI account. Published comments show the display name you choose, not your account email.