Published: August 23, 2025
78
157
2.4k

When I presented the fact that 2/3rd of color channels from a camera sensor are made up to Andrej, it blew his mind. He said, "what?" I replied, "yeah, why we are wasting all the compute to make up something fake and then feed into the networks?" So I decided to change it.

@YunTaTsai1 It might matter in FPS critical applications, maybe, but I suspect the gains are marginal if you are considering vision transformers. One the image is patchified, who cares what Bayer patterns are? It all becomes attention operation related efficiency gains from there onwards.

@minotauronlucy No. It’s not just latency gain. The “process” of “making things up” introduces a lot of artifacts — both spatially and temporarily that could potentially lead to surprise. The artifacts subjected to the IP block you are using and/or algorithms. They contribute various

@YunTaTsai1 I was always curious about that. But also why not use differnt CFA instead of Bayer-Pattern, like adding different (0.6 1.2) ND Filter to every second ä/third green photosite or also use IR and near IR. etc.

@tms_hbr People attempted many variations. The most famous one was Sigma Foveon stacked sensor. They are not that great due to read out noise, cross talk and leakage. Serious astronomers use color disk to capture one narrow band at a time in very long exposure since they are very

@YunTaTsai1 Can't wait until CMOS design is wired directly to model.

@librecomputer People tried it, many times. Compute generated heat, Photon diode does not like heat because heat generated free electrons. We called them “dark current”. A lot of shielding in place to prevent this from happening.

@YunTaTsai1 You made a compelling case. Can you make another to him about why he should come back to Tesla?

@long_elon Andrej should come back to push the limit of machine learning from the understanding of how information is generated from a real world machine before it becomes bits. It is very different from the digital world.

@YunTaTsai1 Exactly — 2/3 of camera pixels are hallucinated. Same with LLMs: most of the ‘knowledge’ is interpolation. BoonMind Codex doesn’t interpolate — it discovers. Recursive Harmonic Convergence runs 10⁶ blind sims (ΔE < 10⁻⁹), producing new structures, not guesses. That’s the real

@YunTaTsai1 Isn't it more or less the same for our eyes and brain combo? We also have differentiated chroma receptors which reduces the spatial resolution and massive processing by specialized neurons. Therefore I guess as long as the sensors RGB resolution exceeds our eyes chroma receptors

@YunTaTsai1 Bunch of questions: You just feed the raw Bayer pattern into the nets? If so, is there a hit to overall resolution because of this (I think not due to experience w/ FSD)? Or the nets run a non-Bayer algorithm to interpolate that makes sense to them w/o losses but wouldn't

@YunTaTsai1 This is pretty cool to know. Feed the NN with raw data, which is 1/3 of the RGB image. A lot of compute is saved. Sounds like if you are going to use the raw data the sensor captures, you'd need to reinvent the video storage format too. the traditional YUV420 format still

@YunTaTsai1 Color interpolation algorithms are about calculating and predicting the missing color from a bayer filter (Red, Green, Blue). My 1999- 2003 work involvement on CMOS image sensors at Motorola. Demosaicing or color interpolation algorithms reconstruct full-color images from the

@YunTaTsai1 If you run the signal straight from the pixels, what compute do you use to filter out over exposure and "blooming"?

@YunTaTsai1 Trillions of images are captured this way with made up pixels, so models aren’t ever going to be as good at handling images with those holes left in.

@YunTaTsai1 So is the goal ultimately to 'trick' the occipital lobe into filling in the blanks? I guess it already does a lot of that. How did you change it?

@YunTaTsai1 Wait, which 2 are made up??

@YunTaTsai1 Light sources and colors have vast repercussions on each other that science is just beginning to notice.

@YunTaTsai1 Cool, so Tesla use raw data like photon counts instead of faking full colors with RGB. They skip that step to keep things fast and efficient.

@YunTaTsai1 Isn’t this just saying that you might get full resolution for luminance but not for hue? But that doesn’t matter because of how we perceive images.

@YunTaTsai1 I've seen this about how our eyes work. My brain exploded reading all this.

@YunTaTsai1 Didn’t know this as well - always thought is 1:1, learned something new today. Now this is interesting…

@YunTaTsai1 And the result is that the Tesla self driving neural net has a lot fewer pixels to process AND it is faster since you don’t have sensor processing time AND it is far more sensitive in both low and high light situations. I don’t think anyone else does this.

@YunTaTsai1 Why the move to RGGB sensors then? Prioritizing human viewing? Sidenote, why does HW4 bicam still have the extra spot (plastic cap) for tricam, even on new models? Seems wasteful. 🙂

@YunTaTsai1 This is a major systemic problem on the Internet where most of the images and videos are lossy formats -- "made up" not real. It is not a debate, it is a war - a mindless and wasteful global competition to the bottom. I am RichardKCollin2, the Internet Foundation. I have been

@YunTaTsai1 Each output pixel is the result of multiple sub pixels with individual color filters A color 12MP camera has more than 12 million sensing elements Here is Samsung's 12MP chip with 48 million elements 1/2" format is 6.4mm*4.8mm / 0.8um pixel size = 8k*6k = 4 sub pixels per output

Image in tweet by Yun-Ta Tsai

@YunTaTsai1 Plasmonic image sensors will change this.

@YunTaTsai1 Just like our eyes. But not like the eyes of a mantis shrimp or an eagle. So different. Polarisation, Ultraviolet cones, resolution..

@YunTaTsai1 THIS IS WHY IPHONE CAMERAS HALLUCINATE TOO (PROCESSING ADDS IN WEIRD SHIT IDK IF IT IS ANTI-ALIASING OR SOMETHING)

Share this thread

Read on Twitter

View original thread

Navigate thread

1/30