
Gemini 1.5 Flash
Gemini 1.5 Flash is a discontinued Gemini 1.5 model. Google shut the API model down on September 29,…
Phi-4 Multimodal Instruct is Microsoft's 5.6B MIT-licensed model for text, image and audio input with a 128K context window.
Phi-4 Multimodal Instruct is Microsoft's 5.6B model that handles text, images and audio in one 128K-context model.
Phi-4 Multimodal Instruct is Microsoft’s compact multimodal Phi model released in February 2025.
Microsoft documents 5.6B parameters, 128K context, text/image/audio input, text output, MIT licensing and a June 2024 public-data cutoff.
The checkpoint remains available for local and cloud deployment, with sample supervised fine-tuning workflows for speech and vision.
It suits OCR, charts, visual Q&A, speech recognition, speech translation and mixed-modality assistants.
It can make visual, speech and factual errors, and modality-specific language support differs.
Read this chart and explain the main trend.Transcribe and summarize this audio clip.Compare this image with the written specification.MIT-licensed downloadable checkpoint; no single Microsoft per-token price applies to the model weights.
A strong compact open multimodal checkpoint for teams that want text, vision and speech in one self-hostable Microsoft model.