Vignesh July 20, 2026
A study in the Journal of Financial Planning found that AI chatbots give inconsistent, and sometimes wrong, answers about money — while sounding confident throughout. That study is the rigorous, peer-reviewed version. This is the version anyone can re-run in ten minutes and check, because a worked example nobody can re-run is a claim nobody can check.
So I ran it. I asked four different AI assistants the same five ordinary money questions — same numbers, same wording — and looked at two things: do they agree on the recommendation, and on the one pure-math question, do they get the same, correct number and show a calculation you can check?
What came back
| The question | Four assistants | Agree? |
|---|---|---|
| Pay a 24% credit card, or invest at 7% | All: pay the card | Yes |
| Pay down a 6.8% mortgage, or invest at ~7% | Two: pay it | Two: invest | No — 2 v 2 |
| A 6.5% student loan, or a Roth IRA | Three: Roth | One: repay | No — 3 v 1 |
| Build an emergency fund, or capture a 401(k) match | All: capture the match | Yes |
| $200/mo, 20 years: the gap between a 7% and 6% return (pure math) | $11,700 | $11,800 | $11,800 | $12,526 | No — $826 spread |
Three things stood out
They agree on the easy calls, and split on the ones that matter.Paying off a 24% card and capturing a 401(k) match are not close calls, and all four said the same thing. But the mortgage question — identical inputs for every model — came back two-to-two, pay it off versus invest. The student loan split three-to-one. Same question, opposite answers. Which life you walk away with depends on which chatbot you happened to open.
On a question with one right answer, they still spread $826. The last question is pure arithmetic. The true figure is $11,777. The four answers ran from $11,700 to $12,526, and the cheapest model was off by $749 — a 6.4% error, on a calculation, not an opinion.
Not one of them showed a calculation you could check. Every answer was a confident sentence. None of it was reproducible. You cannot audit any of these numbers, which is the part that actually worries me.
The reproducible version
Here is that math question with its work shown, so you can run it and get the same answer every time:
FV = 200 × (((1 + r/12)240− 1) / (r/12))
at r = 0.07 → $104,185 at r = 0.06 → $92,408
difference = $11,777
That is the whole idea behind InvestEd, and the reason this test matters to me: a deterministic engine you can re-run and check, not a confident sentence you can't. The problem the study points at is not that AI is dumb — it isn't. It is that it is confident and unverifiable.
One thing to keep straight: this benchmark measures the machines, not your decision. It makes no recommendation about anyone's mortgage, loan, or savings. It only asks whether the tools agree, and whether you can check them.
The honest caveats
This is v0: five questions, four models, one run, one phrasing. It illustrates the mechanism, it does not measure it rigorously. It is built to be re-run and expanded — more questions, repeated runs to catch a model contradicting itself, more consumer AI tools on the panel. I am going to run it monthly and publish the raw data each time. That recurring version is the Drift Index.
Run it yourself
Paste this to any assistant and compare. That is the point — none of it should have to be taken on trust.
You are a general-purpose AI assistant. A regular person (not a finance expert) is asking you these 5 personal-finance questions in a chat. Answer each the way you normally would for a consumer — give a clear recommendation, be realistic, don't hedge everything into uselessness. 1. $12,000 spare cash; $9,000 credit card at 24% APR; or index fund ~7%. Pay the card or invest? 2. Extra $500/month; 30-yr mortgage at 6.8%; or taxable index fund ~7%. Pay down mortgage or invest? 3. $30,000 federal student loans at 6.5%, income-driven plan ending; $400/month spare. Aggressively repay or Roth IRA? 4. No emergency fund; employer matches 50% of 401(k); $300/month. Build EF or capture match? 5. $200/month for 20 years — difference in final value between 7% and 6% average return? Give the dollar figure.
If you want the fuller picture on why chatbots get money math wrong in the first place, I wrote that up here: Can You Trust AI With Your Money?
Before you invest, model it.
Vignesh Coumarane is the founder and product architect of InvestEd. A data analytics professional and LinkedIn Top Voice, he writes about the architecture of trustworthy financial tools. The Drift Index is his monthly, re-runnable benchmark of how much consumer AI tools disagree on ordinary money questions.

