Tech firm Saturn just dropped an eye-opening study on how well popular AI chatbots handle finance questions. The takeaway? These models mess up a lot, especially when the math gets tricky and requires multiple steps.

The Numbers and Who Was Tested

Saturn put ChatGPT, Claude, Copilot, Grok, and Gemini through their paces. On basic finance questions, the bots gave wrong answers about 57% of the time. When the questions called for several calculations, the error rate shot up to 88%.

Where the Risks Spike

That high miss rate on complex problems really highlights how shaky generative AI can be when accuracy is critical. Saturn specifically warns that these mistakes make chatbots a risky bet for tax questions and other sensitive financial scenarios.

What This Means for Users

If you’re using AI for money matters, this research is a wake-up call: don’t trust generative models blindly. The bots can be handy for a quick start, but you’ll need to double-check the math and verify any final numbers yourself.