Hacker Newsnew | past | comments | ask | show | jobs | submit | faizshah's commentslogin

I think I missed why is this faster? What I’m reading here is it’s similar to constrained decoding but I’m not seeing the explanation of why it’s able to get those results.

This thread should be on the first page of every book on microservices.

This or you just repeat the initial prompt every 200k tokens


Which gets you to the point where the whole thing is.. still unreliable. Generative text engines are going to generate. This calls for real enforcement in deterministic pre-edit hooks.

And here is where naive people will say something like "Why do I care if robots shit all over the codebase? Code is for machines, I don't expect to deal with it much now". But really externalized CoT like this confuses machines too, wastes tokens, and eventually wastes exponentially many tokens. Agents tend to think it's more real grounding than prompts are, even for comments-in-code. One bad comment poisons everything, then gets copied around as a ground-truth assumption everywhere. Hooks are more real to them than prompts or comments, and even then if you add enforced limits and tell them to externalize CoT ONLY in scratch task-tracking docs.. they will violate comment-enforcement hooks about 25% of the time. That tells you everything you need to know: even with constant reinforcement, they just really want to break this kind of rule.


Yup. If only there were a task completion hook that could be set to fire prior to rendering terminal output. That would more handily address all these issues, as we could simply enforce output style rules that way.

The current output style does work, but it’s a Sisyphean task to tweak it constantly only to find out that CC adhere’s to only 75% of it, no matter what…


My take on this is they are a tool to help speed up your work they are not meant to produce finished work. Humans produce finished work. LLM will never be deterministic cause their entire value is that they are generalizable.


I hear that, sort of, but here's the thing. Using AI at scale means AI needs to be nearly perfect about not shitting where they eat. That's the subject matter of the whole thread

So the options are a) being a really aggressive stickler for generative hygiene with deterministic rules, b) being massively wasteful about hiring a few machine janitors for every machine coder, or c) humans become the machine's janitor. If I haven't missed an option.. only the first option seems reasonable here.


"AI at scale" is just a euphemism for slop. Current gen AI can augment human engineers but not outright replace them.


This isn't really responsive to what I'm saying or what the discussion here is about, but if you insist. Would you describe lots of AI augmenting lots of human engineers as perhaps.. AI at scale?


I think Elder Scrolls Oblivion was like 80 person core team and 4 years of development and New Vegas was like 70 person dev team and 2 years of development.

But starfield, which was widely criticized by fans, was over 100 people and 8 years development. With the most common criticism from former devs being about the additional management structure and difficulty of communication.

I can’t help but feel there is a lesson there for tech companies where engineering orgs for products are 200+ people when a 20-30 person startup is your top competitor in the vertical.


You can also reduce team size by employing insane, life-destroying crunch time.


My favorite is 20th Century Food Court in Last Call BBS. Some of the other games remind me too much of work (I’ve bought all of them including the coincidence card games but I have only beaten this one), whereas this one reminds me of fun times I had making synths in Logic and VCV Rack for fun. Highly recommend!

I always wished they would make a management or simulation game, I think 90% of all programmers play Paradox games or Tycoon games etc. and I know their take on it would be amazing.


Not really if it takes you 15 minutes to write a 50 line function but it takes the AI 90 seconds then you already are at a 10x speedup just for this task.

This (non-yolo mode AI coding) is actually how we used to code in the old days (2023).


As someone who has used multiple vibe coded internal tools: you will care when you use these tools and encounter strange bugs and missing features.

The human touch is visible in the way your features work just like in vibe coded art and games it lacks intention.


Scale to zero is very useful.


When people say “you’re holding it wrong” it tells me they can’t even conceive of a better way of doing it.

The models produce the same slop for everybody, you don’t have a special way of doing it you lack taste and an opinion on your problem domain from lack of research and studying prior work.


If anyone else couldn’t find the working paper link (the readme and conf link didn’t work for me) it’s this one here: https://github.com/antoinezambelli/forge/blob/main/docs/forg...


Thank you! I've been trying to catch those replies and redirect people, but hopefully your comment be upvoted for others. Very embarrassing to put up the post with the wrong link lol.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: