Gemini has unlocked a new capability: conversational image segmentation š¼ļø This enables new use cases that were previously not possible, furthering Geminiās SOTA image understanding capabilities! š§µ
Read more: https://developers.googleblog.... Great work by @RohanLikesAI @AniBaddepudi and the team!
@OfficialLoganK Is this rolled out?
@CalimanuLoredan yes
@OfficialLoganK Wow! Would be nice if Gemini pro knew how to send a request with thinkingbudget=0 I couldn't find it in docs and Pro gave wrong answers.
@dmytro_petryna No thinking off for pro yet
@OfficialLoganK Remember building this one out last year with API but glad now it's already part of it.
@OfficialLoganK Gemini 2.5: Conversational Image Segmentation I. Executive Summary This briefing document details the advancements in AI visual understanding, specifically focusing on Google's Gemini 2.5 and its new conversational image segmentation capabilities, as outlined in the Google
@OfficialLoganK I posted all https://x.com/GozukaraFurkan/s...
@OfficialLoganK Congratulations team. This was something on radar for long. Will be so helpful to integrate into my application
@OfficialLoganK Interesting š§
@OfficialLoganK Whoa, conversational segmentation is here? Thatās big. Imagine a drone just chatting back while it crops out faulty solar panels or a DAO curator pulling specific layers from an NFT in real time. Howās Gemini holding up with messy, live feedsāany numbers on latency or accuracy
@OfficialLoganK It really is fantastic Logan and Team. I love seeing that youāve pushed. This is really helpful. Really helpful.
@OfficialLoganK is this already available? new model checkpoint? only on flash? also on pro?
@OfficialLoganK Image seg opens doors for richer img analysis.. Think automated labeling & precise object manipulation.. Exciting.
@OfficialLoganK so close
@OfficialLoganK Fail again. Long way to go:
@OfficialLoganK Another fail:
@OfficialLoganK Google is just working without making much noise
@OfficialLoganK This is super cool. Something like understanding a logo would be a big help for founders. Getting the colours right. Or understand the swoop at the end of the design for example. This is a great step up from current models.
@OfficialLoganK these features should all be added to the main chat of the ai studio IMO it doesn't make sense to make every type of output modality its own separate thing you know considering these models are supposed to be "omnimodal" they seem rather sectioned off
@OfficialLoganK nice excited to try it again! - had fun benchmarking detection + seg a few weeks ago and sharing results some interesting notes from the blog, any guidance from the team on: - why flash and not pro as the suggestion? is the new update more fine tuned there - im still shocked
@OfficialLoganK Wasnāt this already available. Gemini models were able to detect object and respond back in JSON coordinates.
@OfficialLoganK The progress thatās been made & importance of Geminiās world understanding is so underrated š
@OfficialLoganK This announcement is much more important than it seems. I don't think it's going to be a function with the marketing and vitality it deserves. This evolution in the understanding of Gemini images is very important for the future development of Multimodal AI that understands
@OfficialLoganK This is really awesome. Many applications
@OfficialLoganK @grok translate what this nerd is trying to say
@OfficialLoganK Looks really cool. Is it just API for now? Will it be in AI Studio?
@OfficialLoganK The 'conversational' part is where the real engineering starts.
@OfficialLoganK masked image > isolate masked area > upscale = all those sf film scenes of "computer, enhance."
@OfficialLoganK For the love of all that is holy, let people use aistudio with their Gemini API key. Getting cut-off mid job is ridiculous.
@OfficialLoganK Cool tech, but are we solving the right problem? Instead of better image segmentation, maybe we need AI that can look at less images and still understand context.
@OfficialLoganK Super impressive!
@OfficialLoganK Great but it's not super super functional yet ā unless 2.5 Pro upgrades that massively. See these: https://x.com/lazharichir/stat... https://x.com/lazharichir/stat... https://x.com/lazharichir/stat...
@OfficialLoganK Can it conversationally parse documents?
@OfficialLoganK If I switch on my camera & share my screen for the whole work day with Gemini access to it. When I switch off my computer after asking it to copy me and log in the next day, I should have Gemini asking me to sit back and watch me working.




