minimaxiH3 Get access
MicrosoftText-to-imageVendor content policy applies

MAI-Image-2.5-Pro Text-to-image

Microsoft AI's highest-fidelity text-to-image model, generating photorealistic, visually dense scenes from natural language with strong object and character consistency and accurate in-image text.

$0.18per imageStarting price at the base resolution and quality tier.
$1.80Ten outputs at the base tier. Larger sizes and premium quality tiers cost more.
Pay per useBilled per request from your balance. No subscription, no minimum.
MAI-Image-2.5-Pro Text-to-image sample output
Sample output · prompt: “A stylish male streetwear model captured in a dynamic mid-air dance pose, one knee lifted high, body slightly twisted, arms extended with powerful movement, wearing a white mesh tank top, oversized black athletic shorts,…”

What it does

MAI-Image-2.5-Pro Text-to-image, in practice.

MAI-Image-2.5-Pro is Microsoft AI's highest-fidelity image model to date — a text-to-image generation model built for work where quality is the top priority. It uses a diffusion-based generative approach that progressively refines the image, producing strong alignment between the input text and the generated output. Compared with the standard MAI-Image-2.5, the Pro tier is tuned for robust object consistency, stronger visual reasoning, and deeper world knowledge, which makes it especially reliable on visually dense, multi-object scenes and on prompts that leave details to be inferred. Released in public preview on July 23, 2026 (model version 2026-06-19), it is the premium tier of the MAI-Im

  • Object consistency across complex scenes — Objects retain the same identity, materials, proportions, markings, and orientation throughout a visually dense composition.
  • Character consistency across views and moments — A person or character remains recognizably the same across poses, camera angles, expressions, clothing, and lighting conditions.
  • Material and physical-property accuracy — Materials look and behave according to their real-world properties, including reflection, translucency, weight, texture, and deformation.
  • Spatial and geometric reasoning — Objects are placed coherently in three-dimensional space, with credible scale, perspective, occlusion, and structural relationships.
  • Photorealistic image synthesis — Realistic imagery with consistent visual structure, natural lighting, accurate skin tones, depth, and texture.
  • High-fidelity portraits — Expressive, natural-looking portraits with accurate facial structure, lighting, and skin texture.

Run MAI-Image-2.5-Pro Text-to-image

from $0.18/image
Invite code opens Chat with this model loaded. No code yet? Join the waitlist — we count which models people ask for.

Parameters

What you can set.

ParameterWhat it doesOptions
promptText prompt describing the image to generate. Maximum context length: 32,000 tokens.
sizeImage dimensions in width*height format (e.g., 1024*1024, 1280*720). Minimum 768. The product of width × height must not exceed 1,048,576.default 1024*1024

Sample prompt

The prompt behind the sample.

A stylish male streetwear model captured in a dynamic mid-air dance pose, one knee lifted high, body slightly twisted, arms extended with powerful movement, wearing a white mesh tank top, oversized black athletic shorts, white crew socks, white sneakers, black gloves, futuristic blue sunglasses and a silver chain necklace. Athletic physique, braided hair, confident and energetic attitude. Minimalist cool-gray studio background, soft diffused lighting, subtle shadows, clean editorial composition. High-end sportswear campaign, contemporary street fashion, cinematic fashion photography, dynamic c
size: 1360*768enable_base64_output: enable_sync_mode:

FAQ

Short answers.

How much does MAI-Image-2.5-Pro Text-to-image cost?

Pricing starts at $0.18 per image at the base resolution; higher resolutions and longer durations cost more. Usage is billed per request from your balance — no subscription.

Does MAI-Image-2.5-Pro Text-to-image run uncensored here?

No. Microsoft applies its own content policy to this model regardless of where it is called from. For a route without an extra platform filter use Uncensored MiniMax H3; this page lists MAI-Image-2.5-Pro Text-to-image for everything else it does well.

What does MAI-Image-2.5-Pro Text-to-image take as input?

It is a text-to-image model. Microsoft AI's highest-fidelity text-to-image model, generating photorealistic, visually dense scenes from natural language with strong object and character consistency and accurate in-image text.

How do I use it?

Enter an invite code to open Chat with the model loaded, or join the waitlist. We open seats in batches and track which models are requested most.

Related

Models people compare with this one.

Run MAI-Image-2.5-Pro Text-to-image.

Invite code opens Chat with the model loaded. No code — join the waitlist and we will count the request.

Get access Browse models