Published: May 21, 2025
4
27
91

10 notable AI models of the week: ▪️ Aya Vision ▪️ INTELLECT-2 ▪️ MiniMax-Speech ▪️ SWE-1 ▪️ Seed1.5-VL ▪️ BLIP3-o ▪️ Skywork-VL ▪️ Behind Maya ▪️ MiMo ▪️ AM-Thinking-v1 🧵

Image in tweet by TuringPost
Image in tweet by TuringPost

1. Aya Vision by @cohere Scales multilingual multimodal generation using synthetic annotation and cross-modal merging, outperforming larger models in multilingual visual reasoning tasks. https://arxiv.org/abs/2505.087... Models Aya-Vision-8B: https://huggingface.co/CohereL... Aya-Vision-32B: https://huggingface.co/CohereL...

Image in tweet by TuringPost

2. INTELLECT-2 by Prime Intellect Team Trains a reasoning model via globally decentralized reinforcement learning, innovating with asynchronous infrastructure and outperforming the 32B reasoning baseline. https://arxiv.org/abs/2505.072...

Image in tweet by TuringPost

3. MiniMax-Speech by @MiniMax__AI Achieves zero-shot high-quality speech synthesis with a learnable speaker encoder and Flow-VAE, excelling in cloning, emotion control, and multilingual support. https://arxiv.org/abs/2505.079...

4. SWE-1 by @windsurf_ai Accelerate software engineering workflows by integrating tool-aware reasoning, incomplete state tracking, and timeline-based context for superior agentic coding experiences. https://windsurf.com/blog/wind...

5. Seed1.5-VL by @ByteDanceOSS Advances agent-centric multimodal reasoning with a compact vision encoder and MoE-based LLM that outperforms competitors on 38 of 60 benchmarks and leads in GUI and gameplay tasks. https://arxiv.org/abs/2505.070...

Image in tweet by TuringPost

6. BLIP3-o by @salesforce Unifies image understanding and generation by using a diffusion transformer for CLIP feature generation and sequential training that preserves comprehension while enabling generation. https://arxiv.org/abs/2505.095... Code: https://github.com/JiuhaiChen/... Models: https://huggingface.co/BLIP3o/...

Image in tweet by TuringPost

7. Skywork-VL Reward by @Skywork_ai Provides multimodal reward signals using a preference-aligned model that improves mixed optimization techniques and sets new standards on VL-RewardBench. https://arxiv.org/abs/2505.072... Model: https://huggingface.co/Skywork...

8. Behind Maya Supports cultural and linguistic diversity in VLMs with multilingual image-text pretraining across eight languages to improve low-resource language understanding. https://arxiv.org/abs/2505.089... Code: https://github.com/nahidalam/m...

Image in tweet by TuringPost

9. MiMo by Xiaomi Unlocks structured reasoning by combining large-scale pretraining with reinforcement post-training on math and coding, beating much larger models on reasoning tasks. https://arxiv.org/abs/2505.076... The model checkpoints: https://github.com/xiaomimimo/...

Image in tweet by TuringPost

10. AM-Thinking-v1 by Beike Elevates open-source reasoning to the 32B scale by fine-tuning Qwen2.5 using RL and SFT, resulting in state-of-the-art scores in math and code. https://arxiv.org/abs/2505.083...

Image in tweet by TuringPost

Read other interesting AI/ML news of the week, as well as freshest research, in our free newsletter: https://www.turingpost.com/p/f...

Share this thread

Read on Twitter

View original thread

Navigate thread

1/12