Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Can any LLM give you the rough pixel coordinates of an item it identifies in an image?

I found that while Claude, GPT etc could describe an image, there was no way to link the description back to specific pixels in the image itself. Not even to a bounding box or segment.



Not as smart as modern frontier models, but Moondream and Molmo can do that sort of thing.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: