Chapitre 3/4 — Multimodal
•
Introduction
Multimodal
Chargement du contenu...
Contenu du chapitre
1
Leçon 1: Panorama du multimodal
2
Leçon 2: CLIP : aligner texte et image
3
Leçon 3: Recherche d'images par texte avec CLIP
4
Leçon 4: BLIP : captioning d'images
5
Leçon 5: LLaVA et les VLMs : question answering visuel
6
Leçon 6: GPT-4V, Claude Vision, Gemini : les modèles commerciaux
7
Leçon 7: Audio : Whisper pour l'ASR
8
Leçon 8: Génération audio : MusicGen et Bark
9
Leçon 9: Text-to-speech et cloning vocal
10
Leçon 10: TP : captioning d'images avec BLIP
11
Leçon 11: TP : pipeline vidéo → transcription → résumé → image
12
Leçon 12: Synthèse multimodal
Ressources
Aucune ressource
Aucune ressource n'est disponible pour ce contenu.