Hacker Newsnew | past | comments | ask | show | jobs | submit | afro88's commentslogin

IMO a much better test would be designs that aren't AI to begin with. Much more useful to see how well a model can html an image design without slopping it up

> Only an encrypted blind relay to allow for shared editing. The relay doesn't see any of the data.

Would love to know more about how this works then? Is it more or less encrypted P2P?


This doc goes into more detail -

https://github.com/nyblnet/bento/blob/main/docs/security.md

But basically when your bento HTML first gets saved, it generates a number of keys used for managing the permissions to join the Cloudflare DO. You can then send out an invite deck with a key that allow other people to join (with their own generated keys). With this architecture, you can revoke people's access, as well as giving them only read-access to the sharing session for example.


Is this a quote from a book? Beautifully written

Thanks afro88. I was inspired by the storm in Terence's story, and my time on my sailboat when it felt like the sea was getting to be too much. Glad to share.

It brought me down to tears, thanks.

Do you have ko-fi link?


Happy to hear.

I added a ko-fi link to my about page, thanks!


When did that happen with Codex? I thought that was a Claude Code thing

Check the github issues brother, search for "usage". Read through the noise. Look at the dates. Don't forget the closed ones.

It's been happening every week since 5.2 was fresh. Not everyone is effected every time. Sometimes it's regional, other times it's per platform, or per version, with a feature (new or old) either enabled or disabled, or some combination of these things. Sometimes you get resets that nobody else does, other times you don't get the ones that were announced. Sometimes you get emails telling you that boosts on limits you didn't even know you had are expiring. At points it did balance out when they really fucked up metering and were underbilling by absurd factors, but that feels more like being toyed with and experimented on than a genuine mistake.

That someone had the 'brilliant' idea to add banked resets (which do expire, so you have to use or lose them, hope you didn't have other plans) says to me that they are no longer as confident about actually being able to fix this as they once were. I'm sure some amount of it is a function of compute availability and reliability, but the rest are definitely issues with the app and it's a rake they keep stomping on.

Compared to just about any other dev centric service I have paid at least $200/mo for, this does not feel very professional to me. They got $1200 outta me and it never really improved. At points I felt like I was getting my moneys worth, and I did burn through a few hundred million tokens, but the frustration of having your projects and plans interrupted and having to wait is not awesome... and then I used Deepseek V4 Pro and felt sick to my stomach with buyers remorse as I watched what it did with just $10.


It's not about figuring out if it's LLM written though. The style is hard to read and annoying. With the kind of sentences GP was talking about it's actually harder to get the substance.


Curious whether you were just bare asking it questions, or whether you provided it with lessons one by one with instruction that the lesson is the baseline truth etc


This has been the case since the early days. Aider had a bunch of code to be very forgiving with formatting of tool calls (file editing in particular at first). It's just the nature of the beast. It surprises me that Pi doesn't have a lot of this kind of stuff built in too


Maybe I'm too optimistic, but given appropriate skills and references (not just for writing but also reviewing) and intelligent use of subagents for isolated reviews and checks, you can lengthen the leash a bit.

But you still need to properly review plans and PRs to keep a good mental model of the codebase. This effectively limits the number of tasks being done in parallel to maybe 2-3. Though you'll be mentally exhausted and probably start to make mistakes or take shortcuts in reviews yourself.


Similar result on our kotlin coding benchmark at work. It measures how close agents can get to a small mergable PR (according to my team). 20 tasks of varying difficulty, with 5 attempts each, LLM as judge to evaluate accuracy (same outcome and quality but allowing for acceptable variances).

Fable 5 sits ahead of Opus 4.7, but behind Opus 4.6, Sonnet 4.6, Opus 4.8, GPT-5.4, GPT-5.5.

Fable isn't a good coding workhorse. That doesn't mean it's not good for actually complex problems and long horizon tasks (big POCs, complex research and such). But I only have vibes and Anthropics own benchmarks and marketing to guide me there.


I'm starting a repository of LLM reviews [1] with the goal of creating a catalog that is more task-oriented and less marketing-y than corporate blogs or benchmark leaderboards. You seem to have a lot of experience across a bunch of different models: if you have a chance and feel like sharing, you'd be one of the first.

[1] - https://model.reviews/ - all the user-submitted content is CC licensed and will be available for download in periodic dumps.


Does your team then manually decide the results by going over the PRs? I suppose you know what you're looking for now, but isn't this still quite painful?


We selected PRs (real ones we merged over the 6 months prior) and have an "LLM as judge" score how close the AI generated code is to the PR. Same as how other benchmarks do it, but it's with tasks we actually do and code we have decided is actually up to scratch for us


I'd love to read about the predictions that have been wrong (genuinely)


> AI's biggest critic has lost the plot

https://www.theargumentmag.com/p/ais-biggest-critic-has-lost...


He has said every month for the past three years that there is no technical progress left to be made in LLMs and that there is no more room in the market for inference spending to grow.

Here[0] is a fun selection of excerpts from his July 2024 post "How Does OpenAI Survive?"[1]

"I see no signs that the transformer-based architecture can do significantly more than it currently does."

[0]: https://xcancel.com/pathsnotchosen/status/206360940100129633...

[1]: https://www.wheresyoured.at/to-serve-altman/


He’s started a cult around anti AI basically.


That statement isn't proof of anything.

If I start a climate crisis cult, doesn't mean my predictions are wrong


I don’t have much time to compile his predictions. I’ll let someone else do that.

That said, I just wanted to point out the grift he is doing. He’s making money by telling some people what they want to hear.

If AI invents a cure for cancer, he will tell you it’s still useless because it didn’t invent immortality.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: