10 notable AI models of the week: ▪️ Aya Vision ▪️ INTELLECT-2 ▪️ MiniMax-Speech ▪️ SWE-1 ▪️ Seed1.5-VL ▪️ BLIP3-o ▪️ Skywork-VL ▪️ Behind Maya ▪️ MiMo ▪️ AM-Thinking-v1 🧵
1. Aya Vision by @cohere Scales multilingual multimodal generation using synthetic annotation and cross-modal merging, outperforming larger models in multilingual visual reasoning tasks. https://arxiv.org/abs/2505.087... Models Aya-Vision-8B: https://huggingface.co/CohereL... Aya-Vision-32B: https://huggingface.co/CohereL...
2. INTELLECT-2 by Prime Intellect Team Trains a reasoning model via globally decentralized reinforcement learning, innovating with asynchronous infrastructure and outperforming the 32B reasoning baseline. https://arxiv.org/abs/2505.072...
3. MiniMax-Speech by @MiniMax__AI Achieves zero-shot high-quality speech synthesis with a learnable speaker encoder and Flow-VAE, excelling in cloning, emotion control, and multilingual support. https://arxiv.org/abs/2505.079...
4. SWE-1 by @windsurf_ai Accelerate software engineering workflows by integrating tool-aware reasoning, incomplete state tracking, and timeline-based context for superior agentic coding experiences. https://windsurf.com/blog/wind...
5. Seed1.5-VL by @ByteDanceOSS Advances agent-centric multimodal reasoning with a compact vision encoder and MoE-based LLM that outperforms competitors on 38 of 60 benchmarks and leads in GUI and gameplay tasks. https://arxiv.org/abs/2505.070...
6. BLIP3-o by @salesforce Unifies image understanding and generation by using a diffusion transformer for CLIP feature generation and sequential training that preserves comprehension while enabling generation. https://arxiv.org/abs/2505.095... Code: https://github.com/JiuhaiChen/... Models: https://huggingface.co/BLIP3o/...
7. Skywork-VL Reward by @Skywork_ai Provides multimodal reward signals using a preference-aligned model that improves mixed optimization techniques and sets new standards on VL-RewardBench. https://arxiv.org/abs/2505.072... Model: https://huggingface.co/Skywork...
8. Behind Maya Supports cultural and linguistic diversity in VLMs with multilingual image-text pretraining across eight languages to improve low-resource language understanding. https://arxiv.org/abs/2505.089... Code: https://github.com/nahidalam/m...
9. MiMo by Xiaomi Unlocks structured reasoning by combining large-scale pretraining with reinforcement post-training on math and coding, beating much larger models on reasoning tasks. https://arxiv.org/abs/2505.076... The model checkpoints: https://github.com/xiaomimimo/...
10. AM-Thinking-v1 by Beike Elevates open-source reasoning to the 32B scale by fine-tuning Qwen2.5 using RL and SFT, resulting in state-of-the-art scores in math and code. https://arxiv.org/abs/2505.083...
Read other interesting AI/ML news of the week, as well as freshest research, in our free newsletter: https://www.turingpost.com/p/f...









