Most AI pilots do not fail on the model. They fail at the first system that is not allowed to be wrong.
The vendor was probably not lying. The demo really did work. It ran against a copy, with clean data, and nothing downstream of it.
Production is different. There is a system of record, a nightly batch, a month end close, and somebody whose job depends on the number being right. An agent that is right 95 times out of 100 is not a 95 percent win against a general ledger, it is a reconciliation problem you now own forever.
This is not an API problem either. Two systems can both speak clean JSON and still disagree about what an order is, when it becomes final, and who is allowed to change it after that. The disagreement is the integration. The transport was never the hard part.
The test I would apply before signing anything: six weeks after go live, can somebody explain why one specific record holds the value it holds, without opening the model. If the answer is no, you do not have automation. You have something you cannot defend to an auditor.
I have connected 12+ procurement platforms, run 300+ branded storefronts in production, and deployed enterprise SSO 15+ times. Nothing on that list was designed to talk to anything else on it, and most of it is not allowed to go down.
Thirty years of that is why I am not worried about whether a model can write the code. It can. It still does not own the cutover.
If you are buying AI work, ask who owns the cutover and what the audit trail looks like on day one. Not day ninety.