Seems like we've forgotten something already? LLMs are bad with the butterfly effect by their very nature. An extremely slight phrasing change can absolutely trigger a deterministic output decision that is out of touch with reality. Tuning only helps if there is training from a previous input. As far as I know (casual research), there still isn't a great solution to this, only marketing hype.