multimodal ai startups

Launching Next has discovered 36 multimodal ai startups.

WanAI.cloud — AI Video Generator
WanAI.cloud — AI Video Generator

Multimodal AI video creation. Make creativity within reach.

Baseer
Baseer

A vision-language model that outperforms GPT-5 on Arabic OCR

DodoForm
DodoForm

Turn talking, pics, or scribbles into clean, structured data

Doing
Doing

Voice and visual context for AI builders. No subscription.

SoulTalk: AI Video Call
SoulTalk: AI Video Call

Video call your AI companion - actually talk, not just type

Uni-1 by Luma
Uni-1 by Luma

A unified foundation model that thinks in pixels

Qwen3.5 Small
Qwen3.5 Small

0.8B-9B native multimodal w/ more intelligence, less compute

Beni AI
Beni AI

Create your AI Companion and facetime anywhere

Kimi K2.5
Kimi K2.5

Native multimodal model with self-directed agent swarms

GetThis
GetThis

Generate tasks from voice, text, or screenshots.

SAM Audio
SAM Audio

Segment any sound with text, visual, or time prompts

Wan 2.6
Wan 2.6

The next era of multimodal AI for creators is here

NexaSDK for Mobile
NexaSDK for Mobile

Easiest solution to deploy multimodal AI to mobile

Inference Engine by GMI Cloud
Inference Engine by GMI Cloud

Fast multimodal-native inference at scale

TwelveLabs Marengo 3.0
TwelveLabs Marengo 3.0

The most powerful embedding model for video understanding

Google Nano Banana Pro
Google Nano Banana Pro

Gemini 3 Pro Image Generator based on Google Nano Banana

Marble by World Labs
Marble by World Labs

A frontier multimodal world model

Qwen-Image
Qwen-Image

Stunning images and perfect text

Wan 2.2
Wan 2.2

The first open MoE model for AI video generation

Deep Research and Imagen in Le Chat
Deep Research and Imagen in Le Chat

Le Chat dives deep (and fun!)

Don't miss the next Unicorn.

Join 5,000+ founders who get a 5-minute briefing on the top startups shaping the future. Every Friday.