Multimodal
A model that can process and generate multiple types of data, such as text and images.
Multimodal models accept inputs beyond text — images, audio, video, or documents — and integrate them into a unified representation. GPT-4o, Gemini, and Claude 3 are multimodal: you can send an image alongside text and the model reasons across both. Multimodal inference typically costs more than text-only due to the additional tokens consumed by vision encoding.
Termes Associés
Un réseau de neurones entraîné sur de grands volumes de texte pour générer du contenu.
L'unité de base de texte que les modèles de langage traitent et facturent.
A large pre-trained model that serves as the base for many downstream applications.
Eğitilmiş bir yapay zeka modelinin yeni girdiler için çıktı üretme süreci.