The Reliability Test:
The Reliability Test:
Can You Trust an AI to Match Your Transactions?
Can You Trust an AI to Match Your Transactions?
Can You Trust an AI to Match Your Transactions?
We tested 3 frontier AI models by running 8 real reconciliation datasets to see if they can accurately and consistently close the books. Almost every run gave a different answer from the last. This report shares our key discoveries.
We tested 3 frontier AI models by running 8 real reconciliation datasets to see if they can accurately and consistently close the books. Almost every run gave a different answer from the last. This report shares our key discoveries.
This report is for financial leaders deciding whether a general-purpose AI agent is ready to post directly to their ledger and what to check beforehand.
This report is for financial leaders deciding whether a general-purpose AI agent is ready to post directly to their ledger and what to check beforehand.



Method
Stacks tested 8 real reconciliation datasets, ranging from 100 to about 20,000 records, to match transactions across a bank statement against a general ledger. Each dataset was run ten times through three general-purpose AI agents, the Stacks Reconciliation Engine, and two rule-based baselines.
No system was given the correct answer in advance. Every output from each system was scored against the same human-reviewed reference: a set of matches confirmed correct by a person, independent of any of the systems tested.
Running the same task ten times exposed a repeatability gap.
Key findings
On the hardest dataset, one AI agent’s wrong matches ranged from 4 to 98 across ten identical runs on the same dataset. Another agent ranged from 7 to 302. While the inputs did not change, the answers did.
In 22 of 24 cases, the AI agents did not give the same answer twice across ten identical runs. For a process like transaction matching, reliability and consistency are critical.
AI agents found 107 of the 117 matches a rule-based system misses entirely, at least once. However, they could not consistently and reliably find the same matches on a second try.
When AI agents created a false match, it could still balance the books to the cent. The total looks correct even when the underlying pairing was incorrect, so a human reviewer may miss it.
Considering the importance of accuracy in the month-end close, finance teams need assurance that any tool they use will not make critical mistakes.
This white paper investigates our findings to understand how frontier AI model agents compare with rule-based baselines and Stacks Reconciliation Engine, and what financial teams can do about it.
Why Does This Study Matter?
Consider the following scenario. A finance leader uses a general-purpose AI agent to implement an AI-assisted reconciliation process. For several months, the process appears to be working well: matches continue to clear, and the close is completed on schedule.
Then, during an audit, one auto-applied transaction match turns out to be incorrect. It was not a match that was flagged for review and missed, but one the system confidently applied on its own.
How did that happen, and, more importantly, is it likely to happen again next month, or was this simply an isolated occurrence?
This paper investigates that exact question.
Discover how reliable general-purpose AI agents are compared to Stacks Reconciliation Agents and rule-based systems when running the same tasks ten times on identical datasets.
Method
Stacks tested 8 real reconciliation datasets, ranging from 100 to about 20,000 records, to match transactions across a bank statement against a general ledger. Each dataset was run ten times through three general-purpose AI agents, the Stacks Reconciliation Engine, and two rule-based baselines.
No system was given the correct answer in advance. Every output from each system was scored against the same human-reviewed reference: a set of matches confirmed correct by a person, independent of any of the systems tested.
Running the same task ten times exposed a repeatability gap.
Key findings
On the hardest dataset, one AI agent’s wrong matches ranged from 4 to 98 across ten identical runs on the same dataset. Another agent ranged from 7 to 302. While the inputs did not change, the answers did.
In 22 of 24 cases, the AI agents did not give the same answer twice across ten identical runs. For a process like transaction matching, reliability and consistency are critical.
AI agents found 107 of the 117 matches a rule-based system misses entirely, at least once. However, they could not consistently and reliably find the same matches on a second try.
When AI agents created a false match, it could still balance the books to the cent. The total looks correct even when the underlying pairing was incorrect, so a human reviewer may miss it.
Considering the importance of accuracy in the month-end close, finance teams need assurance that any tool they use will not make critical mistakes.
This white paper investigates our findings to understand how frontier AI model agents compare with rule-based baselines and Stacks Reconciliation Engine, and what financial teams can do about it.
Why Does This Study Matter?
Consider the following scenario. A finance leader uses a general-purpose AI agent to implement an AI-assisted reconciliation process. For several months, the process appears to be working well: matches continue to clear, and the close is completed on schedule.
Then, during an audit, one auto-applied transaction match turns out to be incorrect. It was not a match that was flagged for review and missed, but one the system confidently applied on its own.
How did that happen, and, more importantly, is it likely to happen again next month, or was this simply an isolated occurrence?
This paper investigates that exact question.
Discover how reliable general-purpose AI agents are compared to Stacks Reconciliation Agents and rule-based systems when running the same tasks ten times on identical datasets.
Method
Stacks tested 8 real reconciliation datasets, ranging from 100 to about 20,000 records, to match transactions across a bank statement against a general ledger. Each dataset was run ten times through three general-purpose AI agents, the Stacks Reconciliation Engine, and two rule-based baselines.
No system was given the correct answer in advance. Every output from each system was scored against the same human-reviewed reference: a set of matches confirmed correct by a person, independent of any of the systems tested.
Running the same task ten times exposed a repeatability gap.
Key findings
On the hardest dataset, one AI agent’s wrong matches ranged from 4 to 98 across ten identical runs on the same dataset. Another agent ranged from 7 to 302. While the inputs did not change, the answers did.
In 22 of 24 cases, the AI agents did not give the same answer twice across ten identical runs. For a process like transaction matching, reliability and consistency are critical.
AI agents found 107 of the 117 matches a rule-based system misses entirely, at least once. However, they could not consistently and reliably find the same matches on a second try.
When AI agents created a false match, it could still balance the books to the cent. The total looks correct even when the underlying pairing was incorrect, so a human reviewer may miss it.
Considering the importance of accuracy in the month-end close, finance teams need assurance that any tool they use will not make critical mistakes.
This white paper investigates our findings to understand how frontier AI model agents compare with rule-based baselines and Stacks Reconciliation Engine, and what financial teams can do about it.
Why Does This Study Matter?
Consider the following scenario. A finance leader uses a general-purpose AI agent to implement an AI-assisted reconciliation process. For several months, the process appears to be working well: matches continue to clear, and the close is completed on schedule.
Then, during an audit, one auto-applied transaction match turns out to be incorrect. It was not a match that was flagged for review and missed, but one the system confidently applied on its own.
How did that happen, and, more importantly, is it likely to happen again next month, or was this simply an isolated occurrence?
This paper investigates that exact question.
Discover how reliable general-purpose AI agents are compared to Stacks Reconciliation Agents and rule-based systems when running the same tasks ten times on identical datasets.
