Launching Next has discovered 36 multimodal ai startups.
Multimodal AI video creation. Make creativity within reach.
A vision-language model that outperforms GPT-5 on Arabic OCR
Turn talking, pics, or scribbles into clean, structured data
Voice and visual context for AI builders. No subscription.
Video call your AI companion - actually talk, not just type
A unified foundation model that thinks in pixels
0.8B-9B native multimodal w/ more intelligence, less compute
Create your AI Companion and facetime anywhere
Native multimodal model with self-directed agent swarms
Generate tasks from voice, text, or screenshots.
Segment any sound with text, visual, or time prompts
The next era of multimodal AI for creators is here
Easiest solution to deploy multimodal AI to mobile
Fast multimodal-native inference at scale
The most powerful embedding model for video understanding
Gemini 3 Pro Image Generator based on Google Nano Banana
A frontier multimodal world model
Stunning images and perfect text
The first open MoE model for AI video generation
Le Chat dives deep (and fun!)
Find new startups about:
Join 5,000+ founders who get a 5-minute briefing on the top startups shaping the future. Every Friday.
Check your email to confirm your subscription.