Published: September 22, 2025
1
0
0

#Gemini flash image, #NanobananaAI , might be performing "semantic edits" i.e generative image editing at semantic level rather normal generative image editing. It means that that model has image understanding at semantic level for visual elements and concepts between/across

@OpenAI @StabilityAI @bfl_ml @midjourney @runwayml @higgsfield_ai @reve @Kling_ai @Alibaba_Qwen @TencentHunyuan @pika_labs @Hailuo_AI @MiniMax__AI @LTXStudio @vivago_ai @ByteDanceOSS had this intuition, which I latter probed while working on the nano banana hackathon submission/project. Sharing for ref not promotion, code is open source. https://youtu.be/z5Bs9q9jEG4?s...

Vision (Image, Video and World) Models Output What They "Think", Outputs are Visuals while the Synthesis Or Generation (process) is "Thinking" (Reasoning Visually).

Share this thread

Read on Twitter

View original thread

Navigate thread

1/5