There are a bunch of approaches that do this kind of thing to reduce token usage ("semble" came to mind, technically different but functionally similar) but their performance is usually mixed because the models haven't been RL tuned to use them as they have the default tool suite. Combine that with the incentive by Anthropic et al. to make you actually burn through as many tokens as possible and I don't see these kind of things becoming mainstream yet. Maybe once we reach a point where consumers actually care about cost (because LLMs have become commoditized) these cost-reduction approaches become relevant enough to actually finetune the model with them.
Enterprises still have big contracts with github, those companies are imposing tight spending limits now and if the open weight models enable those limits to last a bit longer that's probably quite popular.
Everything is currently pointing towards inference being the main cost driver for LLMs in the future. Test-time-compute requires huge amounts of tokens in inference and makes providing frontier models as services unprofitable.
Anyone not under some kind of export restrictions can scrounge together some GPUs to train a frontier model (hell, even DeepSeek which is under these restrictions could) but providing a service that can compete with OpenAI et al. will prove to be quite costly. 3x improvements in inference are therefore nothing to sneeze at IMO.
Interesting! So did you do any experiments on a relevant subset of the data to test whether LLM performance degrades by introducing a new, presumably unknown to the LLM, format?
Accepting the possibility of committing the old "solving social problems with technological solutions" fallacy:
I wonder if an offtopic channel without history (or only a very limited one) could help here. Something that prevents management from scrolling up to identify the people who post too much. In my company I wouldn't even expect that (surveillance) to happen or have consequences but I always have the possibility in the back of my mind.
The kind of management that would trawl through history to check if I'm being too negative is the kind of management that would never allow this history to vanish.
That doesn't help with the "Who's listening?" problem. In an office kitchen you know exactly who is listening and can tailor your comments appropriately. In an online chat, it's not clear who is lurking or who will log on in 5 minutes and read back.
(Welcome to the site btw!)
reply