Black Friday Recharge Offer, offer ends on November 30
?

MiMo-V2.5

Test provider Q1
Coming soon
?
Text Generation

MiMo-V2.5

mimo-v2.5
Text modelComing soontext-to-text

MiMo-V2.5 is Xiaomi's native full-modal model. It achieves professional-grade agent performance at about half the cost of inference, while outperforming MiMo-V2-Omni in multimodal perception in image and video understanding tasks.

From
Coming soon
View model
?

MiMo-V2.5-Pro

Test provider Q1
Coming soon
?
Text Generation

MiMo-V2.5-Pro

mimo-v2.5-pro
Text modelComing soontext-to-text

MiMo-V2.5-Pro is Xiaomi's flagship model, excelling in general-purpose agent capabilities and complex software engineering.

From
Coming soon
View model
?

mimo-v2-omni

Test provider Q1
Popular
?
Text Generation

mimo-v2-omni

mimo-v2-omni
Text modelPopulartext-to-textimage-to-textvideo-to-textspeech-to-text

MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It combines strong multimodal perception with agentic capability - visual grounding, multi-step planning, tool use, and code execution - making it well-suited for complex real-world tasks that span modalities. 256K context window.

From
$0.4/1M tokens
View model
X

mimo-v2-pro

Test provider Q1
Popular
X
Text Generation

mimo-v2-pro

mimo-v2-pro
Text modelPopulartext-to-text

MiMo-V2-Pro is Xiaomi's flagship foundation model, featuring over 1T total parameters and a 1M context length, deeply optimized for agentic scenarios. It is highly adaptable to general agent frameworks like OpenClaw. It ranks among the global top tier in the standard PinchBench and ClawBench benchmarks, with perceived performance approaching that of Opus 4.6. MiMo-V2-Pro is designed to serve as the brain of agent systems, orchestrating complex workflows, driving production engineering tasks, and delivering results reliably.

From
$1/1M tokens
View model