Microsoft has officially launched two new proprietary AI models — MAI-Image-2.5-Pro for image generation and MAI-Voice-2-Flash for speech processing. Already integrated into products like Bing and Dynamics 365, these models demonstrate up to 89% GPU cost reduction compared to equivalent OpenAI solutions. This marks Microsoft's strategic move to reduce third-party AI dependency.
At a glance: Key facts
- MAI-Image-2.5-Pro delivers enhanced detail in image generation
- MAI-Voice-2-Flash cuts voice processing costs by 32% for call centers
- Dynamics 365 Contact Center shows 89% GPU cost savings
- OneDrive reports 26% increase in file retention after switching to MAI-Image-2.5-Pro
- Models currently power Bing, PowerPoint, Excel and GitHub Copilot
- Microsoft employs 'hill-climbing' methodology for performance optimization
- 50% reduction in speech recognition errors for Dragon Copilot medical transcriptions
- MAI-Code-1-Flash shows 10% better code acceptance rate than Claude Haiku 4.5
Microsoft's new AI models explained
MAI-Image-2.5-Pro is Microsoft's flagship image generation model specializing in detailed visualization and accurate text rendering within images. It ranked second on Arena, the generative media evaluation platform. Technical features include:
- 8K resolution support with preserved detail
- Enhanced semantic object segmentation
- 92% text rendering accuracy in Microsoft's tests
MAI-Voice-2-Flash handles mass-scale speech processing for call centers and voice agents. The model operates twice as fast as its predecessor at $15 per million characters. Key advantages:
- 58 language and accent support
- <200ms latency for 95% of requests
- Automatic background noise adaptation
Cost efficiency and performance
Dynamics 365 Contact Center (used by T-Mobile and EasyJet) achieved 89% GPU cost reduction with MAI-Voice-2-Flash. OneDrive saw 26% increased file retention and 25% lower latency after adopting MAI-Image-2.5-Pro.
Optimization techniques
Microsoft's 'hill-climbing' methodology continuously improves models using real user data. MAI-Code-1-Flash for GitHub Copilot shows 10% better performance than OpenAI and Anthropic alternatives. Benchmark results:
| Model | Acceptance Rate | Tokens per task | Retention Rate |
|---|---|---|---|
| MAI-Code-1-Flash | 72% | 142 | 89% |
| GPT-5.4 Mini | 65% | 158 | 83% |
| Claude Haiku 4.5 | 62% | 161 | 78% |
Current implementations
- Bing Image Creator processes 4.2M daily requests with MAI-Image-2.5-Pro
- Dynamics 365 Contact Center handles 12M monthly sessions with MAI-Voice-2-Flash
- Dragon Copilot serves 28M quarterly patients with MAI-Transcribe-1.5
- Excel and PowerPoint automate 87% routine tasks using new models
- Azure Voice Live offers MAI-Voice-2-Flash as developer service
Comparison with OpenAI
While maintaining OpenAI partnerships for complex tasks, Microsoft's new solutions target specific scenarios for greater cost efficiency.
| Metric | MAI-Image-2.5-Pro | GPT-Image-2 (OpenAI) |
|---|---|---|
| Cost (per million tokens) | $5 (text), $8 (image) | 84% more expensive |
| Latency (P95) | 25% lower | Higher |
| GPU efficiency | 2.5x better | Lower |
| Language support | 32 core languages | 45+ |
| Max resolution | 8K | 4K |
Development roadmap
Microsoft plans expanded deployment across Azure and GitHub Copilot, with industry-specific adaptations for healthcare and finance. Expected updates:
- MAI-Voice integration in Teams for auto-transcription
- Arabic and Hindi support in MAI-Transcribe by late 2024
- Optimization for Nvidia Blackwell GPUs
- Enterprise adoption partner program
Questions & answers
How much cheaper are Microsoft's new models vs OpenAI?
MAI-Image-2.5-Pro reduces GPU costs by 84% versus GPT-Image-2, while MAI-Voice-2-Flash delivers 89% savings in Dynamics 365 Contact Center. Excel automation with MAI-Code-1-Flash saves $1.2M annually per 10,000 users.
Which Microsoft products already use MAI-Image-2.5-Pro?
The model powers Bing Image Creator (4.2M requests/day), PowerPoint Designer, OneDrive auto-retouching, and Excel data visualization.
How does MAI-Voice-2-Flash improve call centers?
Processing voice requests twice as fast with <200ms latency, critical for high-volume calls. In 1M session tests, recognition accuracy reached 94.7% vs competitors' 91.2%.
Why is Microsoft transitioning to proprietary AI models?
The company aims to reduce costs (projected $220M annual savings) and gain technology control. Proprietary models avoid licensing constraints and enable better product-specific customization.
Will these models be available to third-party developers?
While currently focused on internal integration, Microsoft plans MAI-Voice-2-Flash as standalone Azure AI service in Q1 2025. Enterprise partners may gain early access.
What's the energy impact of these new models?
Architecture optimizations reduced energy consumption by 37% for MAI-Image-2.5-Pro and 52% for MAI-Voice-2-Flash versus comparable models at equal performance levels.