Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This link just goes to their website. Last I looked at this project, I was happy that it existed but I was disappointed (given my over-optimistic expectations) for two reasons: 1) It's for the BLOOM model which isn't great compared to somewhat recent gpts. Like I think I read that it's worse than the openai models on a per parameter basis. 2) It's faster than using RAM/SSD as faux VRAM but 'only' by 10x. That was even before LLaMA or its improvements had existed for running locally. So by my old understanding, bloom/petals wouldn't be even as good as those ones even though it technically has more parameters. I wonder are these interpretations still true (assuming they ever were true lol), or did something happen where bloom/petals is much better than that now?

Edit: The petals/bloom publication that I read for the information I put above was https://arxiv.org/abs/2209.01188 published to arxiv on September 2 2022.



A Petals dev here. Recent models indeed outperform BLOOM with less parameters (for English). However, the largest LLaMA still doesn't fit into one consumer-grade GPU, and these models still benefit from increasing the number of parameters. So we believe that the Petals-like approach is useful for the newer models as well.

We have guides for adding other models to Petals in the repo. One of our contributors is working on adding the largest LLaMA right now. I doubt that we can host LLaMA in the public swarm due to its license, but there's a chance that we'll get similar models with more permissive license in future.


> I doubt that we can host LLaMA in the public swarm due to its license

Is there anything in the license that specifically forbids distributed usage? If not, you can run it on Petals and just specify that anyone using it must do so for research purposes (or whatever are the license terms)


It does appear to only support Bloom, which makes it currently useless since there are much better models with fewer parameters that you can run on a single machine.

However, the project has a lot of appeal. Not sure how different architectures will get impacted by network latency but presumably you could turn this into a HuggingFace type library where different models are plug-n-play. The wording of their webpage hints that they’re planning on adding support for other models soon.


> However, the project has a lot of appeal. Not sure how different architectures will get impacted by network latency but presumably you could turn this into a HuggingFace type library where different models are plug-n-play.

I love this "bittorent" style swarms compared to the crypto-phase where everything was pay-to-play. People just sharing resources for the community is what the Internet needs more of.


at some point if you want more resources and have them available with the least latency possible, some sort of pay-to-play market will need to appear

even if the currency is computing resources that you have put into the network before (same is true for bittorrent at scale, but most usage of bittorrent is medium/high latency - which makes the market for low-latency responses not critical in that case)


> at some point if you want more resources and have them available with the least latency possible, some sort of pay-to-play market will need to appear

This already exists, it’s corporations. BitTorrent is free, while AWS S3 - or Netflix ;) - is paid.

OpenAI has a pay to use API while this petals.ml “service” is free.

Corporate interests and capitalism fill the paid-for resource opportunities well. I want individuals on the internet to be altruistic and share things because it’s cool not because they’re getting paid.


AWS, or Google Collab etc resemble more paid on demand cloud instances of something like petals.ml than they resemble Netflix.

I don't see the Netflix model working here, unless they can't somehow own the content rights at least partially. Or, as it happens right now with the likes of OpenAI and Midjourney, they sustain a very obvious long term technical advantage. But long term, it's not clear to me it will be sustainable. Time will tell.


I got worse than 1 token/sec, and yes, wasn't impressed with bloom results, but I believe it's also very foreign language heavy. I haven't tried it yet but I believe flexGen benchmarked faster as well.


A Petals dev here. FlexGen is good at high-throughput inference (generating multiple sequences in parallel). During single-batch inference, it spends more than 5 sec/token in case of GPT-3/BLOOM-sized models.

So, I believe 1 sec/token with Petals is the best you can get for the models of this size, unless you have enough GPUs to fit the entire model into the GPU memory (you'd need 3x A100 or 8x 3090 for the 8-bit quantized model).


Unrelated topic: Your username did not age well, huh ?


https://news.ycombinator.com/newsguidelines.html

Comments should get more thoughtful and substantive, not less, as a topic gets more divisive.

Eschew flamebait. Avoid generic tangents. Omit internet tropes.


Thanks for the guidelines link. I was genuinely not aware of guidelines in the comment section.


After lurking I made this account only to post a joking-not-joking explanation of why Alameda had the weirdly specific credit limit $65,355,999,994 with FTX and why I thought it could be a funny off-by-almost-1000x bug/typo/mishap https://news.ycombinator.com/item?id=34473811 but I think almost no one read my comment because I posted it so late after the thread had scrolled off the front page :(


His account was made ~60 days ago, so I don't think that is the case.


Account created 58 days ago. FTX collapsed in November. So.... Especially likely it was meant to be sarcastic, especially with the "bro" suffix


Do me next.


I think the username is an homage to our zeitgeist.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: