Hacker Newsnew | past | comments | ask | show | jobs | submit | vikramkr's commentslogin

IMO building a harness is not wildly difficult (customize pi?) but the offerings from openai and anthropic are wildly subsidized in the subscriptions so they win by default if you want frontier capabilities. Glm 5.3 flash is great but it's not cheaper than a codex or Claude code 200 dollar sub and it does not have astra or fable level capabilities.

Try 50 lines!

https://minimal-agent.com/

I made my own harness based on this, which I jerry rigged to a Codex sub.


In the openai api and many other rapid you can pretty trivially pin not only the model but also the specific snapshot you want to use by using the model id for that snapshot. It's not that big a deal and has been available for years

Often via api you can pin to the specific snapshot. Yes the model provider can screw you over - it's physically possible for many vendors to screw you over, that doesn't mean that's not a dick move and that you can't expect/push for better behavior

These models have specific behavioral characteristics trained into them from reinforcement learning and prompts optimized for one aren't guaranteed to transfer to the new generation. Think if the difference between gpt 5.4 and 5.5 and then 5.5 to 5.6 for example. 5.5 was "better" than 5.4 for struggled more across compaction boundaries and needed much more precise instructions before 5.6 sol recovered some of 5.4's ergonomics. All from the same lab but each model was trained with specific behavioral patterns that were basically product decisions. I would be quite annoyed to find that a model provider was routing a promt optimized for one model to a different one, especially for a dumber/cheaper non frontier model that's not going to be as good at just figuring out what you meant

> Replicating a paper is just as valuable scientifically as publishing it, but how many careers advance through replication?

A lot. In fields where knowledge is incrementally building on previous work the reason the whole field hasn't collapsed from the replication crisis is that usually the results that are really high impact are replicated in as an initial step in new research building on it. It's almost never the focus of the paper but you'll often find a quick mention in methods/supplemental of some previous work that was verified to be valid by a replication of a key technique etc. you'll have crisis where old tools are found to be problematic and findings end up revisited etc. Plus fields like clinical research where there's an awful lot of focus on replicating findings using staged clinical trials with increasing statistical power to determine if new interventions work - that's driven by regulatory requirements grounded in good science and a lot of people make careers in just that.


Not an expert by any means but the assumption here as I understand it is that the arxiv worthy PDF would not be acceptable or meaningful for impossible to understand proofs. And the lean proof would be meaningless unless the specific expression being proven is human understandable as the direct translation of the question the human is asking in formal form. So proving the negation is not a thing but if you make a subtle mistake in translating the statement you want to prove then obviously the QI is going to be proving the wrong thing. And otherwise you're relying on the correctness of lean as a system and on identifying/preventing if the proof is adversarially exploiting bugs in lean to falsely prove things.

Can't speak for op for but me no - it was just ingrained as something disrespectful to do to books. Sometimes you've got a taboo like eating with your left hand that has a pretty clear reason for coming into existence (you use your left hand to clean up in the bathroom and before hand soap in the era of everyone dying from cholera might as well keep those functions as physically separated as possible) but this is more just something you would be disrespectful to do

I mean yeah. Nuclear waste is much safer than fossil fuel waste - deaths caused, radioactivity, etc

I'm so confused by the point you're trying to make. There's a lot of rhetoric about Georgia and an investigation about how the country list is 1:1 with some mysterious list from 1996 - but like - it's an export control list? Yeah, Washington approved the sale of missiles to Georgia. That's how that list works. Washington has to approve it. Openai is not Washington. They're perfectly reasonably erring on the side of caution and potentially over-complying with export controls. And if they then have to get approval from the feds to export to Georgia - well - our current administration has provided many reasons to over comply with trade related controls and not exactly been a champion of encouraging cross country trade right now. Idk what you expect from OpenAI or why you think it would matter at all that Georgia is a democracy or an ally of the US or is closely aligned with us. Ask Canada and NATO how much that's counted for with this administration.

> the intent is good

What is that supposed to mean? They're trying to be an llm search engine that's not some radical new concept


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: