Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

You would not expect the developers of the model to optimize for a well known benchmark?
 help



Here's 'Generate an SVG of a ring-tailed lemur riding an electric scooter' at reasoning level max: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Quote from the thinking trace:

> I’m thinking about how a helmet would obscure lemur ears, but using an electric scooter helmet seems responsible.

It's pretty solid - face is a little wonky but excellent tail and scooter.


Interesting how the background is almost identical to two of the Astra pelicans'.

What if you asked a front/rear/top-side view of the scene (pelican or lemur)?

https://imgur.com/a/60GrJ0i

Did even better from the front. What's surprising is that it used the exact same colors as the simonw example, despite my prompt only being

> Generate an SVG of a ring-tailed lemur riding an electric scooter. Front view

GPT-6 Astra Extra High


It did that thinking about a helmet but didn't actually include a helmet. And there's no "actually, wait, this is just a cute little image no helmet" thought.

Hmm, the face makes it not work as a one shot artifact.

If you want to see something brutal, ask for a zebra riding a scooter. I have yet to see any models do a credible job at that.

I remember that early image models couldn't generate a cyclops no matter how you prompted it, it would at best put a third eye in the forehead.

It seems like Astra has a color palette it likes.

I would expect it happens more organically where people’s discussion of the benchmark and posted results find their way into the training data, like anything else on the internet.

What is this supposed to mean exactly? Do you think there are developers at OpenAI or Anthropic whose job is to train these state of the art models to draw pelicans riding bicycles? Like how exactly do you expect them to be doing that anyways? Hiring graphic artists to create SVGs of bike riding pelicans and feeding thousands of them into the model's training set?

Generate two images, get a vision model to judge, then RL reward the better one. And yes, OpenAI employees on twitter bragged about the model's SVG capabilities. So there are clearly people working there who care about it.

You gotta look up how RLHF works before you ask a demanding question like this.

Right, but that's not the crazy part. The crazy part is thinking they do it all specifically for pelicans on bikes.

If that was the case, the models would have been producing near perfect outputs for it a year ago.

Instead they are just training on general SVG generation, which in no way should be viewed as "benchmaxxing".


> Do you think there are developers at OpenAI or Anthropic whose job is to train these state of the art models to draw pelicans riding bicycles?

Yes


Do you really think there aren't data annotator services specifically training for SVG drawing? And that those annotators don't have a rich set of frequently requested icons/graphics/etc. that they review and train on?

I'd be shocked if they didn't myself.


How did you decide to go from people whose jobs it is to optimize an LLM for a very specific benchmark that is mostly a fun curiosity at best... to employing the services of data annotation services for a rich set of icons and graphics?

You really have to go out of your way to completely misrepresent what's being claimed here in order to make such a wildly off-topic reply.


There was a HN post that tested this hypothesis a few weeks ago: https://news.ycombinator.com/item?id=49010129

I don't think it's a great rebuttal, quite the opposite actually: it's the same prompt with the most obvioy tweak you can imagine and the results are all the same. That definitely looks like it's been something the models have been trained on, and not some emergent capability (and if you take into account how nonthinking Luna compares to the SotA models from a year an half ago: https://simonwillison.net/tags/pelican-riding-a-bicycle/?pag... it becomes even more obvious)

How about you come up with something yourself instead of just moving the goalpost?

No goalpost was harmed in the above comment.

It stopped being a valuable benchmark proxy quite a few model versions ago. Simon knows it, so is everyone who’s serious about it

Treat it like a bit as is


It's actually not just valuable, but infinitely valuable. Because it's imaginary, there is no way to do it perfectly, and thus the possibilities are endless, and it tests out both the formulation and execution of ideas.

Starting my SVG of a pelican on a bicycle as a service startup today, invest now for infinite returns!

The rate at which the value accumulates depends on release of new models and admittedly it's a small community, so you may be in for a wait! But I still mean it, it is a small but enduring stream.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: